Malicious software analysis and identification method and device

By feature processing and extraction of the application editing interface call sequence of the code to be detected in the sandbox, and malicious code recognition is combined with the identification network, the problem of malicious behavior in the existing technology cannot be accurately identified, and the efficiency and accuracy of malware detection are improved.

CN120068073APending Publication Date: 2025-05-30INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510140584.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art cannot accurately identify malicious behavior of malware, resulting in low detection efficiency and poor accuracy.

Method used

By extracting data from the code to be detected running in the sandbox, determining the application editing interface call sequence, and using the multi-headed attention graph network and multi-dimensional word embedding method for feature processing and extraction, combined with the recognition network for malicious code recognition.

Benefits of technology

It effectively reduces the amount of data to be identified, improves the recognition efficiency and accuracy of malware, and improves the robustness and accuracy of malware detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068073A_ABST
    Figure CN120068073A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious software analysis and identification method and device. The method comprises the steps of performing data extraction on a to-be-detected code running in a sandbox, and determining an application program editing interface calling sequence; performing feature processing on the application program editing interface calling sequence through a pre-constructed multi-head attention map network, and determining a feature extraction map; performing feature extraction on the application program editing interface calling sequence through a preset multi-dimensional word embedding method, and determining a target semantic chain embedding vector; and malicious code identification is carried out according to the feature extraction graph and the target semantic chain embedding vector through a pre-constructed identification network, and a target identification result is determined. According to the method, the common features of the malicious behaviors can be captured by fusing the features of different levels, and accurate classification and detection of the malicious software behaviors are realized, so that the robustness and accuracy of malicious software detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to a method and device for analyzing and identifying malicious software. Background Art

[0002] With the rapid development of computer and Internet technologies, security incidents occur frequently. Among them, large-scale malicious software attacks are the main means of current malicious software intrusion. During the operation of a computer, malicious software attacks the computer system, network, and data, causing computer data loss and damage, and can also access the privacy data of the computer without authorization and steal sensitive data. When some malicious software runs, it also occupies the resources of the computer, resulting in a decline in the performance of the computer system, and can also use the computer as a springboard to attack other computers or networks, thus expanding the scope of security threats. To defend against malicious software intrusion, a computer usually installs system security software, and dynamically detects the actual behaviors of malicious software during execution through the system security software dynamic detection technology, including API (Application Programming Interface) calls, network traffic, modifications to the file system and registry, etc., in order to comprehensively analyze the behavior patterns of malicious software; however, as malicious software developers continuously improve anti-detection means, the attack and hiding means of malicious software are becoming increasingly complex and diverse. In order to avoid detection, attackers often intersperse a large number of redundant API calls with normal behaviors before and after the API sequence with real malicious behaviors, so as to achieve the purpose of confusion and hide their truly malicious operations. These redundant API calls will cause existing malicious software detection methods to be unable to accurately identify malicious software. Summary of the Invention

[0003] The present invention provides a method and device for analyzing and identifying malicious software to solve the technical problem in the prior art that the malicious behaviors of malicious software cannot be accurately identified.

[0004] According to one aspect of the present invention, there is provided a method for analyzing and identifying malicious software, including:

[0005] Performing data extraction on the code to be detected running in a sandbox, and determining an application programming interface call sequence;

[0006] Performing feature processing on the application programming interface call sequence through a pre-constructed multi-head attention graph network to determine a feature extraction graph;

[0007] Performing feature extraction on the application programming interface call sequence through a preset multi-dimensional word embedding method to determine a target semantic chain embedding vector;

[0008] Performing malicious code identification based on the feature extraction graph and the target semantic chain embedding vector through a pre-constructed identification network to determine a target identification result.

[0009] According to another aspect of the present invention, there is provided a malicious software analysis and identification device, including:

[0010] A data extraction module, configured to extract data from the code to be detected running in a sandbox to determine an application programming interface call sequence;

[0011] A graph feature module, configured to perform feature processing on the application programming interface call sequence through a pre-constructed multi-head attention graph network to determine a feature extraction graph;

[0012] A semantic feature module, configured to extract features from the application programming interface call sequence through a preset multi-dimensional word embedding method to determine a target semantic chain embedding vector;

[0013] A classification module, configured to perform malicious code identification based on the feature extraction graph and the target semantic chain embedding vector through a pre-constructed identification network to determine a target identification result.

[0014] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the malicious software analysis and identification method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the malicious software analysis and identification method according to any embodiment of the present invention when executed by a processor.

[0019] The technical solution of the embodiment of the present invention extracts data from the code to be detected running in the sandbox to determine the application programming interface call sequence. By extracting the application programming interface call sequence of the code to be detected, the application programming interface call sequence is obtained, which can effectively reduce the amount of data to be recognized and help improve the recognition efficiency of malicious software. The application programming interface call sequence is processed by a pre-constructed multi-head attention graph network to determine a feature extraction graph. The feature extraction graph can effectively reflect the association and importance degree between each call information in the application programming interface call sequence, and can effectively improve the recognition accuracy of malicious software. The application programming interface call sequence is subjected to feature extraction by a preset multi-dimensional word embedding method to determine a target semantic chain embedding vector. The semantic feature can effectively reflect the relevance between application programming interface call sequences, help identify the implicit connections in the application programming interface call sequence, and improve the recognition accuracy of malicious software. The pre-constructed recognition network performs malicious code recognition according to the feature extraction graph and the target semantic chain embedding vector to determine a target recognition result. By fusing features at different levels, the common features of malicious behaviors are captured, realizing the accurate classification and detection of malicious software behaviors, solving the technical problem in the prior art that the malicious behaviors of malicious software cannot be accurately recognized, and improving the robustness and accuracy of malicious software detection.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a malicious software analysis and recognition method provided by an embodiment of the present invention;

[0023] Figure 2 It is a flowchart of another malicious software analysis and recognition method provided by an embodiment of the present invention;

[0024] Figure 3 It is a flowchart of another malicious software analysis and recognition method provided by an embodiment of the present invention;

[0025] Figure 4Schematic structural diagram of a malware analysis and identification device provided by an embodiment of the present invention;

[0026] Figure 5 Schematic structural diagram of an electronic device that can be used to implement the embodiments of the present invention is shown. Detailed implementation manners

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] Figure 1 A flowchart of a malware analysis and identification method is provided for an embodiment of the present invention. This embodiment is applicable to the situation of checking and identifying the malicious code of malware. This method can be executed by a malware analysis and identification device, which can be implemented in the form of hardware and / or software, and the malware analysis and identification device can be configured in an electronic device. As Figure 1 shown, the method includes:

[0030] S110. Extract data from the code to be detected running in the sandbox, and determine the application programming interface call sequence.

[0031] Among them, the sandbox can be a security mechanism of a computer for isolating the code to be detected during operation. In the sandbox, the code to be detected can access and modify the resources inside the sandbox without affecting the computer outside the sandbox. Exemplarily, the sandbox can be selected as Windows Sandbox and / or Sandboxie, etc., and the present invention does not limit this.

[0032] Among them, the code to be detected can be a segment of code in a computer application program or a complete computer application program.

[0033] Optionally, the application program editing interface call sequence can be an API sequence composed of the APIs of the code to be detected. It should be noted that during the running process of the code to be detected, APIs will be called. The APIs called during the running process of the code to be detected are arranged and concatenated in the actual call order to form the application program editing interface call sequence. Among them, the call order of the APIs is composed of the recorded timestamps of calling the APIs.

[0034] Optionally, in a computer, load a sandbox tool. After the sandbox tool is loaded, start the sandbox, put the code to be detected into the sandbox tool, run the code to be detected in the sandbox tool, monitor the dynamic behavior of the code to be detected in the sandbox tool, and record activity information such as system calls, file operations, network communications, and process creations during the running of the code to be detected. Extract data from the APIs called in the activity information to obtain the application program editing interface call sequence.

[0035] Specifically, run the code to be detected in the sandbox, extract data from the activity information of the code to be detected in the sandbox to obtain the application program editing interface call sequence.

[0036] S120: Perform feature processing on the application program editing interface call sequence through a pre-constructed multi-head attention graph network to determine a feature extraction graph.

[0037] Among them, the multi-head attention graph network can be a graph neural network pre-constructed for performing feature processing on graph-structured data. It should be noted that the multi-head attention graph network can be a network structure obtained by combining the multi-head attention mechanism and the graph network. The multi-head attention graph network can process the features of nodes and edges in the graph network through the multi-head attention mechanism, can better capture the complex relationships in the graph-structured data, and can effectively identify and process the node relationships in the graph-structured data.

[0038] Optionally, when the multi-head attention graph network processes the application program editing interface call sequence, it converts the application program editing interface call sequence into graph-structured data and performs feature processing on the converted graph-structured data through the multi-head attention graph network.

[0039] Among them, the feature extraction graph can be the feature data of the graph data structure obtained by the multi-head attention graph network performing feature processing on the application program editing interface call sequence. It should be noted that the feature extraction graph can clearly show the association relationships between the APIs.

[0040] Specifically, convert the application programming interface (API) call sequence into graph structure data, and perform feature processing on the converted graph structure data through a pre-constructed multi-head attention graph network to determine the feature extraction graph corresponding to the API call sequence of the application.

[0041] S130. Extract features from the API call sequence of the application through a preset multi-dimensional word embedding method to determine the target semantic chain embedding vector.

[0042] Among them, the multi-dimensional word embedding method can map the APIs in the API call sequence of the application to a vector space. Exemplarily, the multi-dimensional word embedding method can be k-word embedding.

[0043] Optionally, the process of the multi-dimensional word embedding method can be: for each API in the API call sequence of the application, map each word of the API to a vector space, and represent each word as a multi-dimensional word vector to achieve the extraction of the semantic features of the API call sequence of the application.

[0044] Among them, the target semantic chain embedding vector can be the feature vector obtained by the multi-dimensional word embedding method for semantic feature extraction of the API call sequence of the application; it should be noted that the multi-dimensional word embedding method can sequentially extract features from the APIs in the API call sequence of the application to obtain the target semantic chain embedding vector corresponding to the API call sequence of the application. The target semantic chain embedding vector can represent the linkage relationship between the various APIs in the API call sequence of the application.

[0045] Specifically, sequentially extract features from each API of the API call sequence of the application through the preset multi-dimensional word embedding method. After all the APIs in the API call sequence of the application are extracted, the target semantic chain embedding vector composed of semantic feature vectors is obtained.

[0046] S140. Perform malicious code recognition through a pre-constructed recognition network according to the feature extraction graph and the target semantic chain embedding vector to determine the target recognition result.

[0047] Among them, the recognition network can be a pre-constructed neural network for malicious code recognition. It should be noted that malicious code recognition includes an encoder-decoder sub-network and a classification network.

[0048] Among them, the target recognition result can be the recognition result of the code to be detected. Exemplarily, the target recognition result can be that the code to be detected is malicious code or normal code.

[0049] Specifically, perform malicious code recognition on the feature extraction graph and the target semantic chain through the encoder-decoder sub-network and the classification network to determine the target recognition result.

[0050] In the technical solution of the embodiment of the present invention, by extracting data from the code to be detected running in the sandbox, the application programming interface call sequence is determined. By extracting the application programming interface call sequence of the code to be detected, the application programming interface call sequence is obtained, which can effectively reduce the amount of data to be recognized and help improve the recognition efficiency of malicious software. By performing feature processing on the application programming interface call sequence through a pre-constructed multi-head attention graph network, a feature extraction graph is determined. Through the feature extraction graph, the association and importance degree between each call information in the application programming interface call sequence can be effectively reflected, and the recognition accuracy of malicious software can be effectively improved. By performing feature extraction on the application programming interface call sequence through a preset multi-dimensional word embedding method, a target semantic chain embedding vector is determined. Through the semantic features, the relevance between application programming interface call sequences can be effectively reflected, which helps to identify the implicit connections in the application programming interface call sequence and improve the recognition accuracy of malicious software. By performing malicious code recognition according to the feature extraction graph and the target semantic chain embedding vector through a pre-constructed recognition network, a target recognition result is determined. By fusing features at different levels, the common features of malicious behaviors are captured, and the accurate classification and detection of malicious software behaviors are realized, solving the technical problem in the prior art that the malicious behaviors of malicious software cannot be accurately recognized, and improving the robustness and accuracy of malicious software detection.

[0051] Figure 2 FIG. is a flowchart of another malicious software analysis and recognition method provided by an embodiment of the present invention. The relationship between this embodiment and the above embodiment is the specific method for performing feature processing on the application programming interface call sequence. As Figure 2 shown, the method includes:

[0052] S210. Extract data from the code to be detected running in the sandbox to determine the application programming interface call sequence.

[0053] Optionally, in another optional embodiment of the invention, the extracting data from the code to be detected running in the sandbox to determine the application programming interface call sequence includes:

[0054] Record the log of the code to be detected running in the sandbox to obtain log behavior data;

[0055] Intercept and record the application programming interface of the log behavior data through a preset intercept record function to obtain the call timestamp corresponding to the application programming interface.

[0056] Perform sequential positioning matching extraction on the application programming interface (API) editing interfaces of the log behavior data to determine at least one API editing interface call record, and determine the call sequence information based on the call timestamps of each of the API editing interface call records; perform data filtering on the API editing interface call records to obtain API editing interface call information; and perform sequential combination on the API editing interface call information according to the call sequence information to determine the API editing interface call sequence.

[0057] Among them, the log behavior data can be log events recording various dynamic behaviors of the code to be detected in a sandbox environment. Optionally, the sandbox tool can automatically record the dynamic behaviors of the code to be detected, record each dynamic behavior in the form of an event timestamp, and store it as log behavior data after the sandbox tool runs. Among them, the log behavior data can be a JSON file.

[0058] Among them, the interception recording function can be a function preset to capture an API when the code to be detected calls the API. Exemplarily, the interception recording function can be a hook function.

[0059] Optionally, inject an interception recording function into the sandbox. When the code to be detected calls the API, capture the call information through the interception recording function to obtain the call information and call timestamp of the API; Exemplarily, the call information of the API includes information such as the API name, parameters, return value, etc.

[0060] Among them, the call timestamp can be the time record information when the code to be detected calls the API.

[0061] Among them, the API editing interface call record can be the record information of the API call made by the code to be detected in the sandbox. It should be noted that in the log behavior data, each log event records the event content and the event timestamp, and there are corresponding log events for the API calls of the code to be detected. Therefore, the log events corresponding to the API calls of the code to be detected are used as the API editing interface call records.

[0062] Among them, the call sequence information can be the sequence information of the API calls made by the code to be detected.

[0063] Optionally, for the log behavior data, the application programming interfaces of the log behavior data are sequentially located, matched, and extracted to determine at least one application programming interface call record, and the application programming interface call records are sorted based on the call timestamps of each application programming interface call record to obtain call sequence information. Exemplarily, when the log behavior information is a JSON file, the data is organized by nested "key-value" pairs. The MD5 value of the application programming interface call record is obtained from the log file according to the path ['target']['file']['md5'], and then the application programming interface call record is obtained according to the path ['behavior']['processes']['calls']['api s'].

[0064] Among them, the application marking interface call information can be the name information of the API called in the application programming interface call record. Optionally, since the application programming interface call record includes the request information and response information of the API call, and the covered information content is too much, which affects the data processing efficiency. Therefore, only the name information of the API in the application programming interface call record is selected.

[0065] Optionally, after obtaining the application programming interface call record, the data in the application programming interface call record is filtered to remove the request information and response information, and only the name of the API is retained as the application programming interface call information. Exemplarily, the format of the application programming interface call information can be "NtAllocateVirtualMemory".

[0066] Optionally, for the application programming interface call information of a code to be detected, each application programming interface call information is sequentially combined according to the call sequence information, and all the application programming interface call information is combined into an application programming interface call sequence. Exemplarily, the application programming interface call sequence can be represented as follows;

[0067] ["NtCreateFile","NtWriteFile","NtCreateSection","NtClose","NtMapViewOfSection"]

[0068] Specifically, log the code to be detected in the sandbox tool. After the code to be detected in the sandbox tool finishes running, obtain the log behavior data of the code to be detected. Sequentially match and extract the APIs of the log behavior data through regular expressions to obtain at least one application programming interface (API) call record for editing, and determine the call order information of the API call record for editing. Filter the data of the API call record for editing to obtain the API call information for editing. Combine the API call information for editing in sequence according to the call order information to determine the API call sequence for editing.

[0069] S220. Construct a one-hot encoding matrix according to the API call sequence for editing the application.

[0070] Among them, the one-hot encoding matrix is a matrix form obtained by performing one-hot encoding on the API call sequence for editing the application.

[0071] Optionally, after extracting the API call sequence for editing the application, construct a node set for the API call information for editing in the API call sequence for editing the application. Use each API call information for editing as a node. Assign a unique index to the nodes in the node set for one-hot encoding, and perform one-hot encoding. Arrange all the one-hot vectors in rows to construct a one-hot encoding matrix. Exemplarily, the node set is represented by V, and the node set is expressed as V = {v1, v2,..., Vn}, where n is the total number of API call information for editing in the API call sequence for editing the application. Assign a unique index value to each node in the node set V for one-hot encoding, and perform one-hot encoding. Arrange all the one-hot vectors in rows to form an N*N matrix O. Exemplarily, the matrix O can be expressed as:

[0072]

[0073] Among them, each 1 in the matrix O is the index position corresponding to each node, and 0 represents the index positions of other nodes.

[0074] S230. Use the API call information for editing in the API call sequence for editing the application as graph nodes, and construct the edge relationship between the graph nodes according to the call order information.

[0075] Among them, the graph nodes can be nodes obtained by converting the application programming interface (API) call information into a graph data structure; the edge relationships can be the call order and dependency relationships among the API call information. It should be noted that the graph nodes are the basic units in the graph data structure. In the present invention, the API call information is used as the graph nodes, and the order and dependency relationships among the nodes are identified based on the API call information and the call order information.

[0076] Optionally, each data in the graph data structure is represented in the form of nodes and edges. In the present invention, the API call information in the API call sequence is used as the nodes, and each pair of adjacent calls in the API call information is traversed in sequence according to the call order information, and an edge relationship between two nodes is constructed for the adjacent API call information.

[0077] S240. Perform graph conversion on the API call sequence to determine the API call graph.

[0078] Among them, the API call graph can be the graph data converted from the API call sequence. Optionally, in the API call graph, the nodes represent the API call information, and the edges represent the call order information of each API call information.

[0079] Specifically, the API call information in the API call sequence is used as the nodes, and the call order information of the API call sequence is used as the edges between the nodes. Graph conversion is performed on the API call sequence to determine the API call graph. Exemplarily, the API call graph is represented by G, and G=(V, E) represents the nodes and edges of the API call graph. V can represent the node set of the API call graph, and E can represent the edge set of the node set. V={v1, v2,..., vn}, where n is the total number of API call information in the API call sequence. In the present invention, the node set V can be represented as:

[0080] V = {"NtCreateFile", "NtWriteFile", "NtCreateSection", "NtClose", "NtMapViewOfSection"};

[0081] The edge set E of the node set V can be expressed as: E = [NtCreateFile - NtCreateSection - NtMapViewOfSection, NtCreateFile - NtWriteFile - NtClose].

[0082] The graph G can be expressed as:

[0083]

[0084] Among them, 1 indicates that there is an edge between nodes, and 0 indicates that there is no edge between nodes.

[0085] S250. Perform feature processing on the one-hot encoding matrix and the application programming interface call graph through the multi-head attention graph network to determine the feature extraction graph.

[0086] Specifically, perform feature processing on the application programming interface call graph through a pre-constructed multi-head attention graph network to determine the feature extraction graph corresponding to the application programming interface call sequence.

[0087] Optionally, in another optional implementation of the present invention, the performing feature processing on the application programming interface call graph through the multi-head attention graph network to determine the feature extraction graph includes:

[0088] Input the application programming interface call graph into a pre-constructed multi-head attention graph network. For each graph node, calculate the attention coefficient of the graph node as the target central node of the application programming interface call graph through the attention coefficient calculation formula to obtain the attention coefficient of the graph node.

[0089] For each graph node, perform feature processing on the application programming interface call graph through the multi-head attention graph network according to the attention coefficient of the graph node to determine the feature extraction graph.

[0090] Among them, the target central node can be a node in a key position in the application programming interface call graph. It should be noted that during the process of performing feature processing on the application programming interface call graph, each graph node is sequentially used as the target central node to identify the call and dependency relationships between the target central node and multiple neighbor nodes.

[0091] Among them, the attention coefficient can be an index for the multi-head attention graph network to identify the degree of mutual attention between nodes in the application programming interface call graph; the attention coefficient of a node can reflect the tightness or importance degree of the relationship between the node and adjacent nodes.

[0092] Optionally, the application programming interface call graph is input into a pre-constructed multi-head attention graph network. For each graph node in the multi-head attention graph network, the graph node is used as the central node, and the importance of each neighbor node of the central node to the central node is identified. The information of the neighbor nodes of each node is aggregated through the multi-head attention mechanism to obtain different attention coefficients adaptively allocated by the central node to different adjacent nodes. Exemplarily, in the multi-head attention graph network, the calculation formula of the attention coefficient is as follows:

[0093]

[0094] where, α ij represents the attention coefficient between graph node i and graph node j; x i and x j are the feature vectors of graph node i and graph node j respectively; W is the node weight matrix for linearly transforming the node vectors; represents the transpose of the attention weight a in the attention graph network, which is optimized with training and learning; || represents the concatenation operation; ∑ represents the summation operation; exp represents the exponential function; D i represents the set of all neighbor nodes of graph node i in the application programming interface call graph, j represents graph node j among the neighbor nodes of graph node i; LeakyReLU is the activation function.

[0095] Specifically, the application programming interface call graph is input into a pre-constructed multi-head attention graph network. For each graph node, the attention coefficient is calculated by using the graph node as the target central node of the application programming interface call graph through the attention coefficient calculation formula. For each graph node, the multi-head attention graph network processes the features of the application programming interface call graph according to the attention coefficient of the graph node to determine the feature extraction graph.

[0096] Optionally, in another optional embodiment of the present invention, the step of processing the features of the application programming interface call graph according to the attention coefficient of each graph node by the multi-head attention graph network to determine the feature extraction graph includes:

[0097] Performing feature weighted aggregation according to the attention coefficient of each graph node and the application programming interface call graph to obtain the aggregated feature corresponding to each graph node;

[0098] Updating the features of the application programming interface call graph according to each aggregated feature to obtain the feature extraction graph.

[0099] Among them, the aggregated feature can be the feature vector after the graph node is updated. Optionally, after obtaining the attention coefficient corresponding to each graph node, for each graph node, based on the attention coefficient of the graph node and the features of adjacent nodes in the application programming interface call graph, weighted aggregation is performed to obtain the aggregated feature of the graph node.

[0100] Specifically, according to the attention coefficient of each graph node and the application programming interface call graph, feature weighted aggregation is performed to obtain the aggregated feature corresponding to each graph node, and the aggregated feature corresponding to each graph node is used to update the features of the application programming interface call graph. According to the aggregated features, the features of the graph nodes are updated to obtain the feature extraction graph after the application programming interface call graph is updated.

[0101] Exemplarily, through x′ i ∈{x′ 1 ,x′ 2 ,…,x′ N} where x′ i represents the aggregated feature vector of graph node i, x′ i ∈R F′ ,R F′ represents that the dimension of the feature space where the graph node is located is F′, that is, the feature vector x′ i of the node is a feature vector with a dimension of F′; then the formula for feature weighted aggregation of each graph node is as follows:

[0102]

[0103] where x′ i represents the aggregated feature vector of graph node i; α ij represents the attention coefficient between graph node i and graph node j; D i represents the set of all neighbor nodes of graph node i in the application programming interface call graph, including all neighbor nodes directly connected to graph node i, and j represents graph node j among the neighbor nodes of graph node i; x j represents the feature vector of graph node j, which is the input for aggregation; W is the node weight matrix; σ represents the activation function, which is used to introduce non-linear characteristics.

[0104] Corresponding to each graph node, by using the aggregated features output by each single-head attention layer, the embedding of the node is updated by splicing or averaging the results, so as to capture the correlation between nodes from different angles. The formulas for the two methods of splicing and averaging to update the feature representation of the central node based on the multi-head attention mechanism are as follows:

[0105]

[0106]

[0107] Among them, K is the number of heads in the multi-head attention mechanism, that is, the number of attention heads executed in parallel; is the attention weight of graph node i to its neighbor node j in the k-th attention head; W k is the weight matrix of the k-th attention head; x j is the feature vector of graph node j; D i represents the set of all neighbor nodes of graph node i in the application programming interface call graph, including all neighbor nodes directly connected to graph node i; j represents graph node j among the neighbor nodes of graph node i; represents the concatenation operation of the output features of multiple attention heads (from k = 1 to k = K); σ represents the activation function, which is used to introduce non-linear characteristics.

[0108] S260. Extract features from the application programming interface call sequence of the to-be-detected code running in the sandbox to determine the target semantic chain embedding vector.

[0109] S270. Perform malicious code identification according to the feature extraction graph and the target semantic chain embedding vector through a pre-constructed recognition network to determine the target recognition result.

[0110] The technical solution of the embodiment of the present invention extracts data from the to-be-detected code running in the sandbox to determine the application programming interface call sequence. By extracting the application programming interface call sequence of the to-be-detected code, the application programming interface call sequence is obtained, which can effectively reduce the amount of data to be recognized and help improve the recognition efficiency of malicious software; through the pre-constructed multi-head attention graph network, feature processing is performed on the application programming interface call sequence to determine the feature extraction graph. The feature extraction graph can effectively reflect the association and importance degree among various call information in the application programming interface call sequence, and can effectively improve the recognition accuracy of malicious software; through the preset multi-dimensional word embedding method, feature extraction is performed on the application programming interface call sequence to determine the target semantic chain embedding vector. The semantic features can effectively reflect the relevance between application programming interface call sequences, help identify the implicit connections in the application programming interface call sequence, and improve the recognition accuracy of malicious software; through the pre-constructed recognition network, malicious code identification is performed according to the feature extraction graph and the target semantic chain embedding vector to determine the target recognition result. By fusing the features at different levels, the common features of malicious behaviors are captured, and the accurate classification and detection of malicious software behaviors are realized, solving the technical problem that the malicious behaviors of malicious software cannot be accurately identified in the prior art, and improving the robustness and accuracy of malicious software detection.

[0111] Figure 3The flowchart of another malware analysis and recognition method provided by an embodiment of the present invention. The relationship between this embodiment and the above embodiment is a specific method for extracting semantic features from the call sequence of the application programming interface. As Figure 3 shown, the method includes:

[0112] S310. Extract data from the code to be detected running in the sandbox, and determine the call sequence of the application programming interface.

[0113] S320. Construct a one-hot encoding matrix according to the call sequence of the application programming interface.

[0114] S330. Use the application programming interface call information in the call sequence of the application programming interface as graph nodes, and construct the edge relationship between the graph nodes according to the call order information.

[0115] S340. Perform graph transformation on the call sequence of the application programming interface to determine the call graph of the application programming interface.

[0116] S350. Perform feature processing on the one-hot encoding matrix and the call graph of the application programming interface through the multi-head attention graph network to determine the feature extraction graph.

[0117] S360. For each application programming interface call information in the call sequence of the application programming interface, convert the application programming interface call information into a semantic transformation to obtain the target semantic feature corresponding to the application programming interface call information.

[0118] Among them, the target semantic feature may be a semantic five-tuple describing the application programming interface call information. It should be noted that the target semantic feature is composed of semantic five-tuples. The semantic five-tuples include 5 types of features, namely action feature, operation object feature, system impact grouping feature, category feature, and encoding feature; the action feature describes the action behavior specifically executed by the API; the operation object feature describes the object of action specifically executed by the API; the system impact grouping feature describes the logical or functional attribute of the API; the category feature describes the category identifier of the API; the encoding feature describes the coding specification, language feature, and format convention followed by the API during implementation. Among them, the category identifier may be the classification identifier information pre-stored in the sandbox. Exemplarily, the category identifier may be determined by comprehensively considering multiple factors such as the function, application scenario, and call method of the API, and the category identifier may be set to registry registry.

[0119] Exemplarily, a semantic quintuple <action, object, class, category, encoding>, where action represents the action feature of this API function, object represents the operation object feature of the API function, class represents the system impact grouping feature of the API function, category represents the category feature of the API function, and encoding represents the encoding feature of the API.

[0120] Specifically, for each application editing interface call information in the application editing interface call sequence, convert the application editing interface call information into a semantic transformation to obtain the target semantic feature corresponding to the application editing interface call information.

[0121] Optionally, in another alternative embodiment of the present invention, the converting the application editing interface call information into a semantic transformation to obtain the target semantic feature corresponding to the application editing interface call information includes:

[0122] Match the action of the application editing interface call information with a pre-defined application editing interface action library to obtain the action feature of the application editing interface call information;

[0123] Extract the operation object according to the action feature of the application editing interface call information to determine the operation object feature;

[0124] Classify the system impact grouping of the application editing interface call information according to the action feature to determine the system impact grouping feature;

[0125] Classify the category of the application editing interface call information to determine the category feature;

[0126] Determine the encoding feature according to the encoding information of the application editing interface call information;

[0127] Combine the action feature, the operation object feature, the system impact grouping feature, the category feature, and the encoding feature into the target semantic feature corresponding to the application editing interface call information.

[0128] Among them, the pre-defined application editing interface action library can be the core actions of the commonly used APIs in the sandbox; it should be noted that the number of core actions is at least one, and the present invention does not limit the number of core actions.

[0129] Optionally, match the application programming interface (API) call information with a pre-defined action library of the application programming interface. If a core action matching the API call information is found in the action library of the application programming interface, use this core action as the action feature of the API call information. Exemplarily, the API call information is RegQueryValueExW. When performing action matching in the action library of the application programming interface, the core actions in the action library of the application programming interface can be set to [Exit, Queue, Query, Obtain, Read, Map, Recv, Export, Find, Enum, Protect, First, Exec, Make, Free, Move, Hash, Register, Open, Encode, Gen, Load, Is, Lookup, Put, Next, Encryp, Initial ize, Get]. Therefore, the core action matched by RegQueryValueExW is Query, and the action feature is Query.

[0130] Optionally, after obtaining the action feature of the API call information, identify the degree of impact of the action feature of the API call information on system resource operations and determine the operation object feature. Exemplarily, the operation object feature of RegQueryValueExW can be query system resource 'Select', modify system resource 'Update', or have no impact on the system 'Other'.

[0131] Optionally, after obtaining the action feature of the API call information, identify the system impact grouping feature of the API call information. Exemplarily, the system impact grouping feature of RegQueryValueExW is RegValueEx.

[0132] Optionally, after obtaining the action feature of the API call information, classify the API call information according to the action feature to determine the category feature; Exemplarily, the category feature of RegQueryValueExW can be registry.

[0133] Optionally, after obtaining the action feature of the API call information, identify the coding style in the API call information to obtain the coding feature. Exemplarily, the coding feature of RegQueryValueExW is set to M.

[0134] Specifically, after obtaining the action feature, operation object feature, system impact grouping feature, category feature, and coding feature of each application programming interface (API) call information, combine the action feature, operation object feature, system impact grouping feature, category feature, and coding feature of the application programming interface call information into a tuple sequence to obtain the target semantic feature corresponding to the application programming interface call information. Exemplarily, the five-tuple obtained by decomposing RegQueryValueExW is as follows:

[0135] RegQueryValueExW = <Query, RegValueEX, select, registry, M>

[0136] S370. Convert the application programming interface call sequence into a target semantic chain according to each of the target semantic features.

[0137] Among them, the target semantic chain can be obtained by combining multiple target semantic features;

[0138] Optionally, since each application programming interface call information in the application programming interface call sequence is converted into a target semantic feature, for the application programming interface call sequence, according to the position of each application programming interface call information in the application programming interface call sequence, combine each target semantic feature in turn to obtain the target semantic chain corresponding to the application programming interface call sequence. Exemplarily, if q i represents the target semantic feature corresponding to an application programming interface call information, and the semantic chain formed by combining the target semantic features corresponding to n application programming interface call information is represented by Q, then Q can be represented as:

[0139] Q = [q 1 , q 2 ,..., q i ,..., q n

[0140] S380. Extract features from the target semantic chain through the multi-dimensional word embedding method to determine the target semantic chain embedding vector.

[0141] Optionally, extract features from the target semantic chain through the multi-dimensional word embedding method to obtain the target semantic chain embedding vector.

[0142] Optionally, in another alternative embodiment of the present invention, the step of extracting features from the target semantic chain through the multi-dimensional word embedding method to determine the target semantic chain embedding vector includes:

[0143] ​Perform multi-dimensional word embedding on the target semantic chain through the multi-dimensional word embedding method to obtain a multi-dimensional semantic chain embedding vector;

[0144] Perform convolution processing on the multi-dimensional semantic chain embedding vector to determine the target semantic chain embedding vector.

[0145] Among them, the multi-dimensional semantic chain embedding vector can be the semantic chain embedding vector obtained by performing multi-dimensional word embedding on the target semantic chain. It should be noted that the dimension of the multi-dimensional semantic chain embedding vector is the same as the dimension of the multi-dimensional word embedding method. Exemplarily, for the target semantic chain, k-dimensional word embedding is performed. Since the target semantic chain is composed of five-tuple target semantic features, a 5n×k multi-dimensional semantic chain embedding vector is obtained after embedding.

[0146] Optionally, after obtaining the multi-dimensional semantic chain embedding vector, since the target semantic chain is composed of five-tuples, when performing convolution on the multi-dimensional semantic chain embedding vector after embedding, a convolution kernel with a size equal to the dimension of the multi-dimensional semantic chain embedding vector and a stride of 5 is selected for convolution, which can span all dimensions of an API call, and then the target semantic chain embedding vector is obtained. Among them, the target semantic chain embedding vector is to convert semantic information into a vector form that can be processed by a neural network, which is an abstract representation and can reflect the distribution and relationship of semantic information in the application program edit interface call sequence.

[0147] S390. Perform malicious code recognition according to the feature extraction graph and the target semantic chain embedding vector through a pre-constructed recognition network to determine a target recognition result.

[0148] Optionally, in another optional embodiment of the present invention, the performing malicious code recognition according to the feature extraction graph and the target semantic chain embedding vector through a pre-constructed recognition network to determine a target recognition result includes:

[0149] Fuse the target semantic chain embedding vector and the feature extraction graph to obtain a target fusion feature; input the target fusion feature into the recognition network for malicious code recognition to obtain a target detection result.

[0150] Among them, the target fusion feature can be a fusion feature obtained by sequentially splicing the target semantic chain embedding vector and the feature update graph in the feature dimension direction. Exemplarily, the feature update graph can be represented by F graph and the target semantic chain embedding vector can be represented by F semantics and the target fusion feature can be represented by F fusion . The formula for splicing the target semantic chain embedding vector and the feature update graph can be:

[0151] F fusion =Concat(F graph ,Fsemantics )

[0152] Among them, Concat is used to represent concatenation in the order of the feature dimension direction.

[0153] Optionally, the encoder-decoder sub-network is a Transformer neural network based on the encoder-decoder structure, which can capture long-distance dependencies in long API call sequences through the attention mechanism, enabling the network to capture complex patterns and dependencies in features. The fused features first pass through the encoder part stacked by 6 identical layers, and are linearly mapped to the query, key, and value vector spaces. Then, the attention weights are generated by calculating the similarity between the query and all keys, and the value vectors are weighted and summed to capture the dependencies and importance between different positions in the input features. After being processed by the encoder, the encoder output contains global context information and deep feature representations. The feature vectors processed by the encoder pass through the decoder stacked by 6 identical layers. Using the encoder output as the key and value, the decoder calculates the attention weights with the encoder output when generating the output sequence, enabling the decoder to effectively refer to the global feature information in the encoder and ensuring the consistency and relevance between the generated output sequence and the input features. The entire decoding process iterates through the multi-layer attention mechanism, continuously optimizing the generated feature representations, accurately capturing long-distance dependencies and complex feature patterns, and realizing a profound understanding and effective modeling of the implicit relationships in long sequences. The output of the Transformer network with the encoder-decoder structure is used as the input of the classification network composed of a fully connected network, and classification recognition is performed through the classification network to obtain the target detection result.

[0154] Optionally, in the embodiments of the present invention, it is necessary to train the classification network. By presetting multiple application programming interface call sequences as the training data set, and setting the result label information corresponding to the training data set, each application programming interface call sequence in the training data set is subjected to feature extraction to obtain the feature extraction graph and the target semantic chain embedding vector corresponding to the application programming interface call sequence, and then the features are spliced into training features. The classification network of the recognition network is trained with the training features and the result label information to obtain the trained classification network.

[0155] Specifically, the target semantic chain embedding vector and the feature update graph are fused in features to obtain the target fused features, and the target fused features are input into the recognition network for malicious code recognition to obtain the target detection result.

[0156] The technical solution of the embodiment of the present invention extracts data from the code to be detected running in the sandbox to determine the application programming interface call sequence. By extracting the application programming interface call sequence of the code to be detected, the application programming interface call sequence is obtained, which can effectively reduce the amount of data to be recognized and help improve the recognition efficiency of malicious software. The multi-head attention graph network constructed in advance processes the features of the application programming interface call sequence to determine the feature extraction graph. The feature extraction graph can effectively reflect the association and importance degree among various call information in the application programming interface call sequence, and can effectively improve the recognition accuracy of malicious software. The preset multi-dimensional word embedding method is used to extract features from the application programming interface call sequence to determine the target semantic chain embedding vector. The semantic features can effectively reflect the relevance between application programming interface call sequences, help identify the implicit connections in the application programming interface call sequence, and improve the recognition accuracy of malicious software. The pre-constructed recognition network performs malicious code recognition based on the feature extraction graph and the target semantic chain embedding vector to determine the target recognition result. By fusing features at different levels, the common features of malicious behaviors are captured, and the accurate classification and detection of malicious software behaviors are realized, solving the technical problem in the prior art that the malicious behaviors of malicious software cannot be accurately recognized, and improving the robustness and accuracy of malicious software detection.

[0157] Figure 4 FIG. is a schematic structural diagram of a malicious software analysis and recognition device provided by an embodiment of the present invention. As Figure 4 shown, the device includes: a data extraction module 410, a graph feature module 420, a semantic feature module 430, and a classification module 440; wherein,

[0158] The data extraction module 410 is configured to extract data from the code to be detected running in the sandbox to determine the application programming interface call sequence;

[0159] The graph feature module 420 is configured to process the features of the application programming interface call sequence through a pre-constructed multi-head attention graph network to determine the feature extraction graph;

[0160] The semantic feature module 430 is configured to extract features from the application programming interface call sequence through a preset multi-dimensional word embedding method to determine the target semantic chain embedding vector;

[0161] The classification module 440 is configured to perform malicious code recognition based on the feature extraction graph and the target semantic chain embedding vector through a pre-constructed recognition network to determine the target recognition result.

[0162] The technical solution of the embodiment of the present invention extracts data from the code to be detected running in the sandbox, determines the application programming interface call sequence, and extracts the application programming interface call sequence of the code to be detected to obtain the application programming interface call sequence, which can effectively reduce the amount of data to be recognized and help improve the recognition efficiency of malicious software. By using the pre-constructed multi-head attention graph network to process the features of the application programming interface call sequence, a feature extraction graph is determined. The feature extraction graph can effectively reflect the association and importance degree between each call information in the application programming interface call sequence, and can effectively improve the recognition accuracy of malicious software. By using the preset multi-dimensional word embedding method to extract the features of the application programming interface call sequence, a target semantic chain embedding vector is determined. The semantic features can effectively reflect the correlation between the application programming interface call sequences, help identify the implicit connections in the application programming interface call sequences, and improve the recognition accuracy of malicious software. By using the pre-constructed recognition network to identify malicious code based on the feature extraction graph and the target semantic chain embedding vector, a target recognition result is determined. By fusing the features at different levels to capture the common features of malicious behaviors, accurate classification and detection of malicious software behaviors are realized, solving the technical problem in the prior art that the malicious behaviors of malicious software cannot be accurately recognized, and improving the robustness and accuracy of malicious software detection.

[0163] Optionally, the data extraction module is specifically configured to:

[0164] Record the logs of the code to be detected running in the sandbox to obtain log behavior data;

[0165] Intercept and record the application programming interface of the log behavior data through a preset interception record function to obtain the call timestamp corresponding to the application programming interface;

[0166] Locate and match and extract the application programming interfaces of the log behavior data in sequence to determine at least one application programming interface call record, and determine the call sequence information based on the call timestamp of each application programming interface call record;

[0167] Filter the data of the application programming interface call record to obtain the application programming interface call information;

[0168] Combine the application programming interface call information in sequence according to the call sequence information to determine the application programming interface call sequence.

[0169] Optionally, the graph feature module is specifically configured to:

[0170] Construct a one-hot encoding matrix according to the application programming interface call sequence;

[0171] Take the application editing interface call information in the application editing interface call sequence as graph nodes, and construct the edge relationship between graph nodes based on the call order information;

[0172] Perform graph transformation on the application editing interface call sequence to determine the application editing interface call graph;

[0173] Perform feature processing on the one-hot encoding matrix and the application editing interface call graph through the multi-head attention graph network to determine the feature extraction graph.

[0174] Optionally, the graph feature module is further specifically configured to:

[0175] Input the application editing interface call graph into a pre-constructed multi-head attention graph network. For each graph node, calculate the attention coefficient by using the attention coefficient calculation formula with the graph node as the target central node of the application editing interface call graph, and obtain the attention coefficient of the graph node; wherein, the attention coefficient calculation formula is:

[0176]

[0177] α ij represents the attention coefficient between graph node i and graph node j; x i and x j are the feature vectors of graph node i and graph node j respectively; W is the node weight matrix for linearly transforming the node vectors; represents the transpose of the attention weight a in the attention graph network, which is optimized with training learning; || represents the concatenation operation; ∑ represents the summation operation; exp represents the exponential function; D i represents the set of all neighbor nodes of graph node i in the application editing interface call graph, j represents graph node j among the neighbor nodes of graph node i; LeakyReLU is the activation function.

[0178] Perform feature processing on the application editing interface call graph through the multi-head attention graph network according to the attention coefficient of each graph node to determine the feature extraction graph.

[0179] Optionally, the graph feature module is further specifically configured to:

[0180] Perform feature weighted aggregation according to the attention coefficient of each graph node and the application editing interface call graph to obtain the aggregated feature corresponding to each graph node;

[0181] Perform feature update on the application editing interface call graph according to each aggregated feature to obtain the feature extraction graph.

[0182] Optionally, the semantic feature module is specifically configured to:

[0183] For each application editing interface call information in the application editing interface call sequence, convert the application editing interface call information into a semantic transformation to obtain a target semantic feature corresponding to the application editing interface call information;

[0184] Convert the application editing interface call sequence into a target semantic chain according to each of the target semantic features;

[0185] Extract features from the target semantic chain through the multi-dimensional word embedding method to determine the target semantic chain embedding vector.

[0186] Optionally, the semantic feature module is further specifically configured to:

[0187] Match the actions of the application editing interface call information with a pre-defined application editing interface action library to obtain the action features of the application editing interface call information;

[0188] Extract the operation object according to the action features of the application editing interface call information to determine the operation object feature;

[0189] Classify the application editing interface call information into system impact groups according to the action features to determine the system impact group features;

[0190] Classify the application editing interface call information to determine the category features;

[0191] Determine the encoding feature according to the encoding information of the application editing interface call information;

[0192] Combine the action feature, the operation object feature, the system impact group feature, the category feature, and the encoding feature into the target semantic feature corresponding to the application editing interface call information.

[0193] Optionally, the semantic feature module is further specifically configured to:

[0194] Perform multi-dimensional word embedding on the target semantic chain through the multi-dimensional word embedding method to obtain a multi-dimensional semantic chain embedding vector;

[0195] Perform convolution processing on the multi-dimensional semantic chain embedding vector to determine the target semantic chain embedding vector.

[0196] Optionally, the classification module is specifically configured to:

[0197] Embed the target semantic chain embedding vector and the feature extraction graph for feature fusion to obtain a target fusion feature;

[0198] Input the target fusion feature into the recognition network for malicious code recognition to obtain a target detection result.

[0199] The malicious software analysis and recognition device provided by the embodiments of the present invention can execute the malicious software analysis and recognition method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0200] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0201] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0202] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0203] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the malware analysis and identification method.

[0204] In some embodiments, the malware analysis and identification method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the malware analysis and identification method described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the malware analysis and identification method by any other suitable means (e.g., by means of firmware).

[0205] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0206] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the patterns / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0207] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0208] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0209] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0210] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0211] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0212] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the malware analysis and identification method provided in any embodiment of the present invention. The method includes:

[0213] Extract data from the code to be detected running in the sandbox to determine the application programming interface call sequence;

[0214] Perform feature processing on the application programming interface call sequence through a pre-constructed multi-head attention graph network to determine a feature extraction graph;

[0215] Perform feature extraction on the application programming interface call sequence through a preset multi-dimensional word embedding method to determine a target semantic chain embedding vector;

[0216] Malicious code recognition is performed based on the feature extraction graph and the target semantic chain embedding vector through a pre-constructed recognition network to determine a target recognition result. The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0217] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0218] The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0219] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0220] Those of ordinary skill in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program codes executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0221] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.

[0222] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A malware analysis and identification method, characterized in that: include: Extract data from the code to be detected running in the sandbox and determine the application editing interface call sequence; Performing feature processing on the application editing interface call sequence through a pre-built multi-head attention graph network to determine a feature extraction graph; Perform feature extraction on the application editing interface call sequence by using a preset multi-dimensional word embedding method to determine a target semantic chain embedding vector; Malicious code is identified through a pre-built recognition network according to the feature extraction graph and the target semantic chain embedding vector to determine a target recognition result.

2. The method according to claim 1, characterized in that The step of extracting data from the code to be detected running in the sandbox and determining the application editing interface calling sequence includes: Logging the code to be detected running in the sandbox to obtain log behavior data; The application program editing interface of the log behavior data is intercepted and recorded by a preset interception recording function to obtain a call timestamp corresponding to the application program editing interface; Sequentially locate, match and extract the application editing interface of the log behavior data, determine at least one application editing interface call record, and determine call sequence information based on a call timestamp of each application editing interface call record; Filtering the application program editing interface call record to obtain application program editing interface call information; The application editing interface calling information is sequentially combined according to the calling sequence information to determine the application editing interface calling sequence.

3. The method according to claim 2, characterized in that The method of performing feature processing on the application programming interface call sequence through a pre-built multi-head attention graph network to determine a feature extraction graph includes: Constructing a one-hot encoding matrix according to the application programming interface call sequence; Using application editing interface call information in the application editing interface call sequence as a graph node, and constructing edge relationships between the graph nodes based on the call sequence information; Performing graph conversion on the application editing interface call sequence to determine the application editing interface call graph; The one-hot encoding matrix and the application programming interface call graph are feature processed by the multi-head attention graph network to determine the feature extraction graph.

4. The method according to claim 3, characterized in that The performing feature processing on the application editing interface call graph by the multi-head attention graph network to determine the feature extraction graph includes: The application editing interface call graph is input into a pre-built multi-head attention graph network. For each graph node, the graph node is used as the target central node of the application editing interface call graph to calculate the attention coefficient through the attention coefficient calculation formula to obtain the attention coefficient of the graph node; wherein the attention coefficient calculation formula is: Among them, α ij represents the attention coefficient between graph node i and graph node j; x i and x j are the feature vectors of graph node i and graph node j respectively; W is the node weight matrix, which is used to perform linear transformation on the node vector; represents the transpose of the attention weight a in the attention graph network; || represents the concatenation operation; ∑ represents the summation operation; exp represents the exponential function; D i represents the set of all neighbor nodes of graph node i in the application editing interface call graph, j represents graph node j among the neighbor nodes of graph node i; LeakyReLU is the activation function; The multi-head attention graph network performs feature processing on the application editing interface call graph according to the attention coefficient of each graph node to determine the feature extraction graph.

5. The method according to claim 4, characterized in that The step of performing feature processing on the application programming interface call graph according to the attention coefficient of each of the graph nodes through the multi-head attention graph network to determine the feature extraction graph includes: Performing weighted feature aggregation according to the attention coefficient of each of the graph nodes and the application programming interface call graph to obtain an aggregate feature corresponding to each of the graph nodes; The application editing interface call graph is updated with features according to each of the aggregated features to obtain the feature extraction graph.

6. The method according to claim 1, characterized in that The method of extracting features from the application programming interface call sequence by a preset multi-dimensional word embedding method to determine a target semantic chain embedding vector includes: For each application editing interface call information in the application editing interface call sequence, converting the application editing interface call information into a semantic conversion to obtain a target semantic feature corresponding to the application editing interface call information; Converting the application programming interface call sequence into a target semantic chain according to each target semantic feature; The target semantic chain is feature extracted by the multi-dimensional word embedding method to determine the target semantic chain embedding vector.

7. The method according to claim 6, characterized in that The converting the application editing interface call information into a semantic conversion to obtain a target semantic feature corresponding to the application editing interface call information includes: Performing action matching on the application editing interface call information and a predefined application editing interface action library to obtain action features of the application editing interface call information; Extracting the operation object according to the action features of the application editing interface call information and determining the operation object features; Classify the application editing interface call information into system impact groups according to the action characteristics to determine system impact group characteristics; Classifying the application program editing interface call information into categories and determining category characteristics; Determining a coding feature according to the coding information of the application program editing interface calling information; The action feature, the operation object feature, the system impact group feature, the category feature and the coding feature are combined into the target semantic feature corresponding to the application editing interface call information.

8. The method according to claim 7, characterized in that The step of extracting features of the target semantic chain by using the multi-dimensional word embedding method to determine the embedding vector of the target semantic chain includes: Performing multi-dimensional word embedding on the target semantic chain by using the multi-dimensional word embedding method to obtain a multi-dimensional semantic chain embedding vector; The multi-dimensional semantic chain embedding vector is subjected to convolution processing to determine the target semantic chain embedding vector.

9. The method according to claim 1, characterized in that: The method of performing malicious code identification according to the feature extraction graph and the target semantic chain embedding vector through the pre-built identification network to determine the target identification result includes: Performing feature fusion on the target semantic chain embedding vector and the feature extraction graph to obtain a target fusion feature; The target fusion feature is input into the recognition network to perform malicious code recognition to obtain a target detection result.

10. A malware analysis and identification device, characterized in that: include: A data extraction module is used to extract data from the code to be detected running in the sandbox and determine the application editing interface call sequence; A graph feature module, used to perform feature processing on the application editing interface call sequence through a pre-built multi-head attention graph network to determine a feature extraction graph; A semantic feature module, used to extract features from the application editing interface call sequence by a preset multi-dimensional word embedding method, and determine a target semantic chain embedding vector; The classification module is used to identify malicious code according to the feature extraction graph and the target semantic chain embedding vector through a pre-built recognition network to determine the target recognition result.

Citation Information

Cited By

  • Word set acquisition method, intelligent measurement terminal application detection method and device based on word set, equipment and storage medium

    CN120278152A