A JavaScript malicious code detection method based on a prototype dependency graph
By introducing prototype dependency graphs and graph neural networks into JavaScript malicious code detection, the problem of incomplete semantic information extraction in existing detection methods is solved, and more efficient and accurate malicious code detection is achieved.
Patent Information
- Application Number
- CN202410357830.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-03-27
AI Technical Summary
The existing JavaScript malicious code detection method based on static analysis is difficult to detect obfuscated JavaScript malicious code due to inadequate extraction of code semantic information, resulting in inefficient detection.
It is proposed to use a prototype dependency graph as a code representation method, combine the graph neural network to extract node and graph structure information, and process the graph data through pruning algorithm and pre-trained FastText model, and finally realize the detection of malicious code.
Through the combination of prototype dependency graphs and graph neural networks, the semantic information of JavaScript code can be extracted more comprehensively, the accuracy and efficiency of malicious code detection can be improved, and the problem of incomplete semantic information extraction in existing detection methods is effectively solved.
Smart Images

Figure CN118036004B_ABST
Abstract
Description
[0001] This patent relates to a method for code representation of JavaScript code through a prototype dependence graph to detect JavaScript malicious code. Background Art
[0002] In recent years, due to its ease of use and powerful functions in creating interactive dynamic web pages, JavaScript has been widely used as a client-side programming language in the development of Internet websites. However, while this programming language brings convenience to development work, it also causes network security problems related to JavaScript malicious code. Attackers embed malicious code in web pages using the JavaScript language to achieve purposes such as stealing users' private information or property.
[0003] JavaScript malicious code detection technology is a security defense means for detecting malicious code in websites that has received extensive attention currently. The detection technologies based on dynamic analysis and those based on static analysis are two main research directions for JavaScript malicious code detection technology currently. The dynamic analysis method often uses virtual environments such as honeypots and sandboxes to execute JavaScript code, extracts the behavioral characteristics of the code executor, records and analyzes the execution information, and finds out the typical malicious behavioral characteristics contained therein to classify the code. The static analysis method extracts the characteristics of the code to be detected and matches them with the malicious characteristics obtained after model training to search whether the code to be detected contains suspicious keywords or fragments. It also extracts the semantic structures of the source code, key strings, and functions, and then uses machine learning or deep learning to build a detection model.
[0004] Dynamic analysis can more clearly reveal the behavior of malicious code, but the recognizability of the analysis environment causes malicious code to be able to evade detection by checking the environment. In addition, the overhead of dynamic analysis is too high. Therefore, it is not suitable for large-scale analysis. On the other hand, static analysis consumes fewer resources, is simple and fast in detecting malicious code, and is more cost-effective. However, traditional static analysis-based detection methods are difficult to detect obfuscated JavaScript malicious code due to the insufficient extraction of code semantic information. Summary of the Invention
[0005] Aiming at the problem of low detection efficiency caused by the lack of semantic information in static detection methods, this patent proposes a new code representation method, the prototype dependency graph, to fully represent code semantic information. To better extract node and graph structure information in the prototype dependency graph, a graph neural network is used for feature extraction, and finally a new JavaScript malicious code detection method is realized. Prototypes and prototype chains are the basis for object and inheritance implementation in JavaScript. This patent combines existing code representation methods and adds information about prototypes and objects to make up for the deficiencies of existing code representations. Then, a pruning algorithm is used to prune the prototype dependency graph, and a pre-trained FastText model is used to encode the node types and edge types in the graph, converting the text into vectors. Finally, a graph neural network is used to extract node feature information and graph structure information, and the samples are classified according to the vectors output by the neural network.
[0006] The JavaScript prototype dependency graph proposed in this patent aims to represent the prototype information of the code and make up for the semantic loss of existing representation methods. The definition of the prototype dependency graph will be introduced in detail below.
[0007] The prototype dependency graph is a graphical representation that represents all JavaScript prototypes, objects, variables, and tree nodes generated during abstract parsing as nodes and represents their relationships as edges. These edges include the relationships between prototypes and objects, variables (such as Definitions 1 and 2 below), and the relationships between abstract syntax tree nodes and objects, variables (such as Definitions 3 and 4 below). The edge definitions of the prototype dependency graph are as follows:
[0008] Definition 1 Prototype definition uses an edge - an edge from an object node to a prototype node, representing the prototype node corresponding to the object node.
[0009] Definition 2 Prototype property lookup edge - an edge from a prototype node to a variable, representing the property shared by the instances constructed by the constructor on the prototype.
[0010] Definition 3 Variable definition uses an edge - an edge from an abstract syntax tree node to a variable, representing that the abstract syntax tree node defines or uses the variable.
[0011] Definition 4 Object lookup edge - an edge from a variable to an object node, representing that the variable node is an instance of a certain object node.
[0012] First, introduce the edges related to AST nodes - variable definition and usage. When constructing variable definition edges, the construction of related objects and prototypes needs to be considered simultaneously. If the variable is a primitive data type such as number, string, or boolean, the variable will no longer point to an object. If the variable is a reference data type such as function or object, object lookup edges and prototype definition edges need to be further constructed. When constructing variable usage edges, the construction of objects and prototypes does not need to be considered. When the variable is a reference data type, this edge will be used to reproduce property lookup in abstract parsing.
[0013] There can be multiple object lookup edges because multiple variable names can point to the same object. For example, in the code, there are two variables a and b and an object property obj.v that point to the same object. There are three variable nodes and three edges from a, b, and obj.v to the object node in the prototype dependency graph, but there is only one object node representing this object. Therefore, all three object lookup edges will be resolved to the same object node during abstract parsing.
[0014] There can be multiple prototype definition and usage edges because multiple objects can point to the same prototype. For example, in the code, variables a and b are objects constructed by the same constructor, and there are two function calls a.doMalicious() and b.doMalicious(). Both a and b will use the doMalicious() method in the same prototype. Using prototypes can show the real source of object properties or object methods in the code and describe the data calls when actually running the JavaScript code parser.
[0015] Prototype property lookup edges are all the properties shared by the objects constructed by the constructor in the prototype, enabling the retrieval of all property characteristics that are not explicitly defined when the object is defined but are inherited through the prototype and mounted on the prototype, that is, representing the information of the prototype chain.
[0016] The overall structure of the JavaScript malicious code detection method based on the prototype dependency graph described in this patent is as Figure 1 shown and mainly consists of five parts:
[0017] (1) Sample set construction module. The sample set construction module crawls data of benign code through a web crawler, obtains malicious code from existing public datasets, and then divides the obtained malicious code and benign code datasets.
[0018] (2) Prototype Dependency Graph Generation Module. First, convert the sample code into an abstract syntax tree through Esprima, take the JSON file of the abstract syntax tree, and convert it into a Node node tree; second, extract the control flow graph and program dependency graph of the sample code respectively; third, traverse the entire node tree, and add the control information of the control flow graph and the data flow information of the program dependency graph to the corresponding node tree respectively; finally, add the prototype dependency information defined above to the node tree to generate a prototype dependency graph.
[0019] (3) Graph Data Conversion Module. First, prune the prototype dependency graph according to the maximum number of nodes, removing the nodes exceeding the maximum number of nodes and the connected edges; second, use the FastText word embedding model to convert the text features into vector form; finally, construct the input data of the graph neural network based on the obtained vectors and graph connectivity information.
[0020] (4) Graph Neural Network Learning Module. Use a graph attention network that adds feature extraction on edges to perform feature learning on the graph input data constructed in the previous step.
[0021] (5) Prediction and Output Module. Use the mean aggregation function to obtain the graph-level vector representation and input it into the softmax layer to predict the true label of the sample, and display the result.
[0022] The beneficial effects of the present invention compared with the prior art are as follows:
[0023] (1) JavaScript Prototype Dependency Graph Generation Technology. Aiming at the problem that existing abstract representation methods do not model the prototype call relationships in the code, resulting in the loss of features contained in the prototype relationships during malicious code detection and incomplete extraction of code semantic information, this paper proposes a code representation method of prototype dependency graph. The prototype dependency graph is a graph representation method that adds prototype relationship information to the code abstract representation of JavaScript malicious code detection and combines it with other code abstract representation methods.
[0024] (2) Graph Neural Network Based on Prototype Dependency Graph. Traditional JavaScript malicious code detection methods represent the graph as a sequence, which leads to a large amount of information loss about the graph structure. In order to consider both the attribute features of each node and the structure features of the entire graph, this patent uses a graph neural network to extract features from the prototype dependency graph. This method can extract both the node and edge information on the prototype dependency graph and the structure information of the graph. Description of the Drawings
[0025] To more clearly illustrate the technical solutions in the embodiments of this patent, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this patent. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0026] Figure 1 It is the overall architecture diagram of the JavaScript malicious code detection method based on the prototype dependency graph provided by this patent;
[0027] Figure 2 It is the flowchart for generating the prototype dependency graph provided by this patent;
[0028] Figure 3 It is the example code provided by this patent;
[0029] Figure 4 It is the abstract syntax tree of the example code provided by this patent;
[0030] Figure 5 It is the example prototype dependency graph provided by this patent;
[0031] Figure 6 It is the structure diagram of the graph neural network provided by this patent. Specific embodiments
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this patent clearer, the following will clearly and completely describe the technical solutions in the embodiments of this patent in conjunction with the accompanying drawings in the embodiments of this patent. Obviously, the described embodiments are some, rather than all, of the embodiments of this patent. Based on the embodiments in this patent, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this patent.
[0033] In the embodiments of this patent, the process of the JavaScript malicious code detection method based on the prototype dependency graph is as follows:
[0034] Step 1: Construct a prototype dependency graph. The construction process of the prototype dependency graph is as Figure 2 shown. First, the JavaScript code is converted into an abstract syntax tree through the high-performance parser Esprima. After obtaining the JSON file of the abstract syntax tree, it is read, and then a node tree of the Node type is generated according to the data in the JSON file. The definition of the Node class is shown in the following table:
[0035] Attribute Name Type id int type string name string body dict parent list children list Edge Attribute dict
[0036] parent represents the list of parent nodes. children represents the list of child nodes. In edge_attribute, the key is another node that has an edge relationship with the current node, and the value is the attribute of the edge starting from that node, which stores information about the edges in the subsequent control flow graph and prototype dependency graph, including six string - type tags: boolean, e, prototype definition use, variable definition use, object search, and prototype attribute search.
[0037] After converting the abstract syntax tree into Nodes, it is necessary to generate a control flow graph. When constructing the control flow graph, mainly by performing a depth - first traversal of the Node tree, it is judged whether the type of the current Node is IfStatement and SwitchStatement. If so, the corresponding boolean edge relationship is added to the edge_attribute of this Node. If this node has child nodes, the corresponding e edge relationship is added to the child nodes.
[0038] Since the prototype dependency graph includes the construction of the data transfer process from definition to use of data, in actual coding, the construction of the program dependency graph is not done separately, but is placed in the process of generating the prototype dependency graph.
[0039] The generation process of the prototype dependency graph is introduced in detail below. Perform a depth-first traversal on the control flow graph. If the node type is VariableDeclaration or FunctionDeclaration, it means that the node is the definition of a variable or a function. Add the corresponding variable, object, and prototype nodes to the tree, and add a variable definition edge between the Node node and the variable node, that is, add variabledefinition use in the edge_attribute. Add an object search edge between the variable node and the object node, that is, add object search in the edge_attribute. Add a prototype definition edge between the object node and the prototype node, that is, add prototype definitionuse in the edge_attribute. If the node type is Identifier, it means that the node is using a variable. If the name attribute of the node is non-prototype, that is, not a prototype node, create a variable node and add a variable use edge variabledefinitionuse. If the name is prototype, make corresponding changes to the prototype of the variable and add the corresponding prototype attribute search edge prototypeattribute search. By judging the node types and adding nodes and edges during the above-mentioned depth-first traversal of the tree, the final prototype dependency graph can be generated.
[0040] This patent takes Figure 3 the code shown as an example to elaborate on the finally constructed prototype dependency graph. Through the Esprima tool mentioned above, Figure 3 the code can be converted into Figure 4 the abstract syntax tree shown. Figure 4 The labels in Figure 5 indicate that the abstract syntax tree nodes appear more than once in the graph, which is used to distinguish different nodes in the subsequent prototype dependency graph. According to the definition of the prototype dependency graph described in the invention content and the specific generation process described above, the Figure 5 prototype dependency graph shown is constructed. In Figure 4 , the abstract syntax tree nodes are represented as rectangles (the same shape as the abstract syntax tree in
[0041] Step 2: First, prune the graph. First, perform a breadth-first traversal on the graph, and delete the nodes sorted after the maximum node number threshold and the edges connected to them. Then, use the FastText library in Python to obtain the vectors of the nodes and edges in the prototype dependency graph, and convert them into the form of data in the input format of the graph neural network Data=(X, A, R, Y), where X represents the vector of nodes, A represents the adjacency matrix, R represents the vector of edges, and Y is the label of benign code or malicious code.
[0042] Step 3: Use the graph attention network to extract features from the data input in the previous step, and add edge features to the original graph attention network. The graph structure is as Figure 6 shown. Let h i be the feature vector of node v i in the neural network, and e ij be the feature vector of the edge between node v i and node v j in the neural network. Then, the updated feature h' i of node v i can be calculated by the following formula:
[0043]
[0044] where N(i) represents the set of neighbor nodes of node v i , W is the weight matrix for node feature transformation, W e is the weight matrix for edge feature transformation, α ij is the attention coefficient calculated through the attention mechanism, indicating the importance of node v j for node v i . σ is the activation function LeakyReLU.
[0045] The formula for the processed attention coefficient α ij is as follows:
[0046]
[0047] where || represents the concatenation of vectors, is the transpose of the attention weight vector, and softmax j is the softmax function applied to all j∈N(i) for normalizing the attention coefficient.
[0048] The entire graph model consists of two attention layers and a global pooling layer. Finally, softmax is used to classify using the graph-level feature vector. To complete the prediction for each graph, all nodes in the graph need to aggregate and summarize as much information as possible on a single graph. This patent uses the mean aggregation method to aggregate nodes. The aggregation readout formula for the average of all node features for each graph is as follows:
[0049]
[0050] where h G is the feature representation of graph G, V is the set of nodes in graph G, and h v is the feature of node V. Finally, the vector of the graph-level representation is input into the softmax layer to predict the result. This patent uses the cross-entropy function to minimize the loss:
[0051]
[0052]
[0053] where W and b are the weights and biases, and y i represents the true label value.
[0054] Finally, it should be noted that: The above embodiments are only used to illustrate the technical solutions of this patent and are not intended to limit them; although this patent has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this patent.
Claims
1. A JavaScript malicious code detection method based on prototype dependency graph, characterized in that: include: Step S101, obtaining JavaScript code data, using malicious code in the public data set and combining it with benign code crawled by the crawler to form a sample data set; Step S102, constructing a prototype dependency graph, first converting the JavaScript code into an abstract syntax tree through Esprima, then adding control information and data flow information to form a control flow graph and a program dependency graph, and finally adding nodes and edges related to the prototype to the node tree to form a prototype dependency graph; Step S103, prototype dependency graph data conversion, first pruning the graph, pruning the nodes and connected edges that exceed the threshold, and then using FastText to encode the node and edge information in the prototype dependency graph to form corresponding node vectors and edge vectors; Step S104: training the graph neural network, using the graph attention network with edge feature extraction to perform feature learning on the graph input data constructed in the previous step; Step S105, prediction and output, use the mean aggregation function to obtain the graph-level vector representation and input it into the softmax layer to predict the true label of the sample, and output the result.
2. The prototype dependency graph according to claim 1, characterized in that: The prototype dependency graph is a graphical representation that represents all JavaScript prototypes, objects, variables, and abstract syntax trees generated during abstract parsing as nodes and their relationships as edges. These edges include: (1) Prototype definition uses edges - the edges from object nodes to prototype nodes, indicating the prototype nodes corresponding to the object nodes; (2) Prototype attribute lookup edge: an edge from a prototype node to a variable, representing the attributes shared by instances constructed by the constructor on the prototype; (3) Variable definition and use edge: an edge from an abstract syntax tree node to a variable, indicating that the abstract syntax tree node defines or uses a variable; (4) Object lookup edge: an edge from a variable to an object node, indicating that the variable node is an instance of an object node.
3. The method for detecting JavaScript malicious code based on prototype dependency graph according to claim 1, characterized in that: The step S101 includes: The benign code data is crawled by web crawlers, malicious code is obtained from existing public data sets, and then the obtained malicious code and benign code data sets are divided.
4. The method for detecting JavaScript malicious code based on prototype dependency graph according to claim 1, characterized in that: The step S102 includes: Use Esprima to convert the sample code into an abstract syntax tree, take the JSON file of the abstract syntax tree, and parse it into a Node tree; Traverse the Node tree and add the control information of the control flow graph and the data flow information of the program dependency graph to the Node tree; Add the nodes and edges described in right 2 to the Node node tree respectively to build the final prototype dependency graph.
5. The method for detecting JavaScript malicious code based on prototype dependency graph according to claim 1, characterized in that: The step S103 includes: Perform a breadth-first traversal on the constructed prototype dependency graph, identify and remove nodes and their connected edges that are ranked after the maximum node number threshold, so as to simplify the graph structure; The FastText word embedding model is applied to convert the text features of nodes and edges in the prototype dependency graph into node vectors and edge vector representations respectively, enhancing the semantic richness of the features. Combined with the converted vector representation, an input dataset suitable for graph neural network is constructed.
6. The method for detecting JavaScript malicious code based on prototype dependency graph according to claim 1, characterized in that: The steps S104 and S105 include: Use the graph attention network to extract features from the input data. The graph attention network consists of two attention layers and a global pooling layer. The nodes are aggregated by the mean aggregation method to obtain the graph-level feature vector, and then the graph-level feature vector is classified using softmax.
Citation Information
Patent Citations
Webpage filtering method based on program slicing technology
CN103970845A
Malicious webpage detection method based on multi-modal fusion
CN117313094A