Code similarity detection method based on deep learning

By extracting the syntax and semantic features of the code based on deep learning, the problem that code similarity detection methods in the prior art are difficult to parse semantic features, and more efficient and accurate code similarity detection is achieved.

CN120336871AActive Publication Date: 2025-07-18SICHUAN JINGLANG INTELLECTUAL PROPERTY AGENCY CO LTD

Patent Information

Application Number
CN202510393968.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing code similarity detection methods are difficult to fully analyze the semantic characteristics of the code, and traditional methods are prone to misjudgment when dealing with codes with similar syntax structures but different semantics, and the detection efficiency is low.

Method used

Using a deep learning-based method, the syntax features of the code are extracted through the AST-GCN model, combined with the multi-headed self-attention layer to enhance the syntax features, and the semantic features are extracted using the multi-scale semantic features adaptive fusion model MSF, and finally, the code similarity is calculated through the gating mechanism.

Benefits of technology

It improves the accuracy and efficiency of code similarity detection, can better handle codes with similar syntax structures but different semantics, reduces misjudgments, and improves the comprehensiveness and accuracy of code detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336871A_ABST
    Figure CN120336871A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of code similarity detection, and particularly relates to a code similarity detection method based on deep learning, which comprises the following specific steps of: collecting codes, preprocessing the codes, namely deleting corresponding annotation contents, normalizing variable names and function names, unifying code formats, and storing the codes. Any two-by-two combination and code similarity rating are carried out on the processed codes to obtain a triad lt; a code A, a code B and similarity ygt; and the set of all triples forms a code data set. According to the method, the overall similarity of codes is evaluated by calculating the similarity of grammar and semantic features of the codes; in the aspect of extracting code grammar feature vectors, when feature aggregation is carried out on each node in a grammar tree, the depth where the node is located and the features of brother nodes of the node are considered, and feature loss caused when a tree structure is aggregated through a conventional method is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code similarity detection, and particularly to a code similarity detection method based on deep learning. Background Art

[0002] In software development, code similarity detection can be used on the one hand to detect code copying and plagiarism identification, which helps to improve the quality and efficiency of software development, and promotes code innovation and progress; on the other hand, it can help to locate similar code segments, which is beneficial to code reuse and integration, improves the maintainability and reusability of code, and promotes code sharing. Currently, common code similarity detection methods include the following categories:

[0003] (1) Code similarity detection methods based on code text comparison: These methods use character sequence matching algorithms to compare the character sequences of code texts to calculate the similarity of codes. Common character sequence matching algorithms include the Longest Common Subsequence (LCS) algorithm and the String Pattern Matching (KMP) algorithm. These methods are simple to implement and relatively fast in calculation speed, and are effective for codes with simple structures, small differences in variable names and formats.

[0004] (2) Code similarity detection methods based on syntax trees: These methods use compiler front-end tools to convert codes into syntax trees. The syntax tree represents the syntax structure of the code in a tree-like structure, where nodes represent syntax elements (such as expressions, statements, declarations, etc.), and edges represent syntax relationships. These methods compare the syntax trees of two pieces of code and quantify and evaluate the syntactic similarity of the codes by measuring the similarity between the syntax trees. The similarity between two syntax trees can be compared and calculated using tree comparison algorithms. Common tree comparison algorithms include the tree edit distance algorithm and the subtree matching algorithm. The tree edit distance algorithm calculates the minimum number of edit operations required to convert one tree into another tree, while the subtree matching algorithm searches for the same subtree structures in two trees.

[0005] (3) Code similarity detection methods based on graphs: The basic idea of these methods is similar to that of code similarity detection methods based on syntax trees. By analyzing the syntax structure of the code, as well as function call relationships, control dependencies, data flows, etc., a dependency graph of the program is constructed, the nodes in the graph are matched, and the similarity of the dependency graph can be calculated through graph comparison algorithms or graph convolutional neural networks, thereby performing similarity detection on the codes. In this detection idea, the nodes and connection relationships of the graph reflect the syntax of the code, and the data flow can be regarded as an abstract expression of code semantics. Therefore, it has good similarity detection capabilities and may detect similar codes with different text structures but similar functions.

[0006] However, current code similarity detection faces some challenges. Code similarity detection methods based on code text comparison are very sensitive to code format changes (such as indentation, spaces) and variable name modifications, and it is difficult to handle cases where the semantics are the same but the lexical structures are different. When code similarity detection methods based on syntax trees parse the syntax structure of code, they cannot accurately learn the relationships between nodes, so they cannot obtain results that accurately reflect the syntax features of the code. Moreover, this method is not comprehensive enough in extracting code semantics, resulting in misjudgments in the case of similar code syntax structures but different semantics. Building a Program Dependence Graph (PDG) of the program can reflect the semantic features of the code to a certain extent, but it does not fully contain the syntax features of the code, and the cost of building the program dependence graph is very high and the detection efficiency is low.

[0007] Based on the above, the current methods have the following disadvantages:

[0008] (1) Code similarity detection methods based on code text or syntax trees cannot fully parse the semantic features of code. These methods mainly focus on the syntax form of code and do not fully consider the actual semantics of the code;

[0009] (2) When using a graph convolutional model to traverse and update the node features of a syntax tree, the graph convolutional model cannot fully learn the syntax features of the code through feature aggregation of the tree structure, thus reducing the effect of similarity detection;

[0010] (3) The traditional methods for fusing syntax and semantic features usually use fixed-value weighting or direct concatenation, and cannot be adjusted specifically according to the characteristics of different codes, which will lead to important features being weakened and unimportant features being overemphasized, affecting the effect of similarity detection

[0011] Therefore, a code similarity detection method based on deep learning is invented. Summary of the Invention

[0012] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:

[0013] A code similarity detection method based on deep learning, which includes the following specific steps:

[0014] S1: Collect codes, preprocess the codes, and the operations include deleting the corresponding comment content, normalizing the variable names and function names, unifying the code format, combining any two of the processed codes pairwise, and rating the code similarity to obtain a triple <Code A, Code B, similarity y>. The set of all triples constitutes a code data set, and Code A or Code B is set as Code C;

[0015] S2: Parse the code C using a code extraction tool to generate a syntax tree, and traverse the syntax tree to obtain the adjacency matrix of code C and the depth weight matrix d' C , and then convert code C into a feature vector of a specified dimension through Word2Vec to obtain the feature matrix Z C ;

[0016] S3: Input and Z C into the AST-GCN model. AST-GCN performs syntax feature aggregation on the syntax tree of code C to obtain the syntax feature matrix Finally, perform feature enhancement on through the multi-head self-attention layer to obtain the syntax feature vector V of code C gC ;

[0017] S4: Perform word segmentation on code C, and use the word segmentation tool to split code C into several lexical units to obtain a lexical unit sequence X C ;

[0018] S5: Construct a multi-scale semantic feature adaptive fusion model MSF for multi-scale fusion and extraction of code semantic features. After inputting X C into the MSF model, obtain the feature vector V representing the code semantics sC ;

[0019] S6: Use the gating mechanism to adaptively fuse the syntax feature vector V of code C gC and the semantic feature vector V sC to obtain the fusion vector V fusC ;

[0020] S7: Concatenate V fusA and V fusB to obtain the vector V to be measured. Input V t into a three-layer perceptron for similarity comparison, and output the similarity score t Calculate the loss value loss between and the true value, and update the learnable syntax matrix W gC and the learnable semantic matrix W sC ;

[0021] S8: Apply the trained code similarity detection model to the code similarity detection task, calculate the similarity of any actually input code pair <code X, code Y>, and obtain the similarity score of the code pair;

[0022] The specific steps of the above S3 are as follows:

[0023] S31: Construct an AST-GCN model for extracting the syntax feature matrix of code C The AST-GCN model consists of an input layer, a feature update layer, and an output layer; the AST-GCN model fuses the features of each node itself in the syntax tree with the features of its neighbor nodes and sibling nodes to obtain a feature sequence containing local syntax features of the code Finally, a syntax feature matrix is obtained through a linear transformation

[0024] S32: Use the multi-head self-attention mechanism to calculate the attention weights, concatenate the syntax feature sequences output by each attention head, and then obtain the syntax feature vector V through a linear transformation gC .

[0025] As a preferred solution of a code similarity detection method based on deep learning according to the present invention, wherein: the specific steps of S31 are as follows

[0026] S311: Traverse the syntax tree of code C hierarchically to obtain a sibling matrix and a depth weight matrix d' C ;

[0027] S312: Input d' d' C , and Z C into the feature update layer. The feature update layer consists of L convolutional layers. Assume that the syntax tree performs feature update in the (l + 1)-th layer (l ∈ (0, 1,..., L - 1)) of the feature update layer. Then the feature matrix output in its layer is composed of the addition of three parts of features, namely the sibling feature matrix the hierarchical neighbor feature matrix and the feature matrix output by the previous layer of the feature update layer

[0028] S313: Add and to obtain the updated feature matrix in the (l + 1)-th layer of the feature update layer

[0029]

[0030] where ReLU is the activation function

[0031] S314: After the final layer calculation of the feature update layer, output the feature matrix

[0032]

[0033] Among them i ∈ (1, 2, …, n) is the characteristic sequence of each node in Z C in the The updated feature sequence output by the final layer of the feature update layer is linearly transformed to obtain the syntax feature matrix Make a linear transformation to obtain the syntax feature matrix

[0034] As a preferred solution of the code similarity detection method based on deep learning described in the present invention, wherein: the specific steps of S311 are as follows:

[0035] S3111: Traverse the syntax tree hierarchically. For any node r encountered during the traversal, determine the sibling nodes of r according to the adjacency relationship between nodes, thereby obtaining the sibling matrix In All sibling relationships between syntax tree nodes are stored, that is, if node p and node q are sibling relationships, then there is x p,q 、 x p,q = x q,p = 1; otherwise x p,q = x q,p = 0;

[0036] S3112: When traversing the syntax tree, record the depth of each node in the syntax tree in the depth matrix d C in, d C is a diagonal matrix, and the values on its diagonal are the depths of each node; according to the preset maximum critical depth m, adjust the weight coefficient during feature aggregation. For any node i in d C whose depth is greater than m, the value of i is replaced by where α is the scaling coefficient, d i represents the depth of node i; and for any node j whose depth is less than or equal to m, the value of j is replaced by 1. After the above weight coefficient adjustment, the depth weight matrix d' C is finally obtained;

[0037] The specific steps of S312 are as follows:

[0038] S3121: Calculate the sibling feature matrix Use and to perform feature aggregation operations to obtain

[0039]

[0040] Among them is the inverse matrix of the degree matrix, W s is the weight matrix;

[0041] S3122: For nodes deeper in the syntax tree, in order to reduce the impact of the features of these nodes on the overall code similarity, multiply on the left by the depth weight matrix d' C and then perform feature aggregation to obtain the hierarchical neighbor feature matrix

[0042]

[0043] where W n is the weight matrix.

[0044] As a preferred solution of a code similarity detection method based on deep learning according to the present invention, wherein: the specific steps of the S5 are as follows:

[0045] S51: First, input X C into CodeBERT, perform a concat operation on the semantic sequences output by each self-attention head of CodeBERT to obtain a feature vector that preliminarily represents the code semantics, and then obtain the semantic vector v of code C through a linear transformation C ;

[0046] S52: Perform a downsampling operation on v C to obtain a semantically coarse-grained sequence Then input v C and into the bidirectional simple recurrent unit F-BiRSU with an adaptive forgetting factor to obtain the semantic feature sequence H of code C C ;

[0047] S53: Use the self-attention layer to further enhance the semantics of H C so that H C can incorporate information from other positions in the entire sequence and better capture the global semantic information of the code, and finally obtain the semantic feature vector V of code C sC .

[0048] As a preferred solution of a code similarity detection method based on deep learning according to the present invention, wherein: the specific steps of the S52 are as follows:

[0049] S521: Use max pooling to perform a downsampling operation on v C to obtain a semantically coarse-grained sequence For any its calculation formula is as follows:

[0050]

[0051] in, is the semantic vector of code C, k is the size of the maximum pooling window, and n is an integer multiple of k;

[0052] S522: Build F-BiSRU module for v C and Extract semantic feature sequence H C , its F-BiSRU module consists of a forward FSRU network and a reverse FSRU network, each of which contains several FSRU units;

[0053] S523: and Perform concat operation to get And through a linear transformation, the semantic feature sequence H is obtained C ,Right now:

[0054]

[0055] in

[0056] The specific steps of S522 are as follows:

[0057] S5221: First define the gate vector I C , used to control and v C The fusion ratio of the time series input at each moment, I C The calculation formula is as follows:

[0058]

[0059] Where Sigmoid is the activation function, W t is the weight matrix, is the bias vector, Indicates that v C and The time series at time t and Splice and then and According to I C Perform multi-scale information fusion to obtain

[0060]

[0061] Where ⊙ is an element-wise multiplication operation;

[0062] S5222: Forget gate when calculating the forward FSRU network at time t When outputting, introduce the hidden state of the previous moment Combine with and splice them as the updated parameters. The specific calculation formula is as follows:

[0063]

[0064] Where W f is the weight matrix, b f is the bias vector, is the hidden state of the previous moment of the forward FSRU network, is based on statistical characteristic function of variance, θ is a coefficient used to scale the value;

[0065] The reset gate r t of the forward FSRU network at time t, the memory cell c t and the forward hidden state output at time t are calculated as follows:

[0066]

[0067]

[0068]

[0069] The reverse hidden state output by the reverse FSRU network at time t is calculated in the same way as the forward hidden state output by the forward FSRU network at time t, but when calculating the forget gate the hidden state of the previous moment of the reverse FSRU network is introduced

[0070]

[0071] As a preferred solution of a code similarity detection method based on deep learning described in the present invention, wherein: the specific steps of S6 are as follows:

[0072] S61: Design a linear layer composed of a learnable syntax matrix W gC and a bias value b gC to generate the syntax gating vector g gC of code C:

[0073] g gC = ReLU(W gC V gC + b gC )

[0074] Similarly, design a linear layer composed of a learnable semantic matrix W sC and a bias value b sC to generate the semantic gating vector g of code C sC :

[0075] g sC = ReLU(W sC V sC + b sC )

[0076] Calculate the fusion weight α of the syntactic feature vector of code C C and the fusion weight β of the semantic feature vector C :

[0077]

[0078] β C = 1 - α C

[0079] S62: Use α C , β C to perform weighted summation on V gC , V sC to obtain the final fusion vector V of code C fusC :

[0080] V fusC = α C V gC + β C V sC ;

[0081] Among them, since code A or code B is set as code C in S1, V fusC actually represents the fusion vectors V of code A fusA , V of code B fusB at the same time, that is, after S62, V fusA and V fusB have been obtained.

[0082] As a preferred solution of a code similarity detection method based on deep learning described in the present invention, wherein: The specific steps of S7 are as follows:

[0083] S71: Concatenate the fusion vectors V fusA , V fusB to obtain the vector to be measured V t :

[0084]

[0085] where Indicates a splicing operation;

[0086] S72: Input V t into a three - layer perceptron for training, and output a value between 0 and 1 as the similarity score between Code A and Code B

[0087] Z1 = ReLU(W1V t + b1)

[0088] Z2 = ReLU(W2Z1 + b2)

[0089] Z3 = ReLU(W3Z2 + b3)

[0090]

[0091] where W1, W2, and W3 are weight matrices, b1, b2, and b3 are bias vectors, and Z1, Z2, and Z3 are the feature representations output by the first, second, and third layers in the three - layer perceptron respectively;

[0092] S73: Calculate the loss value loss between

[0093]

[0094] and the true similarity value. The calculation formula is as follows: i where y

[0095] is the true similarity value between Code A and Code B in this training. The closer its value is to 1, the more similar Code A and Code B are; gC and the learnable semantic matrix W sC for back - propagation update.

[0096] As a preferred solution of a code similarity detection method based on deep learning according to the present invention, wherein: The specific steps of S74 are as follows:

[0097] S741: The update process of W gC is as follows:

[0098]

[0099] S742: The update process of W sC is as follows:

[0100]

[0101] where represents the partial derivative, and η1, η2 are the learning rates;

[0102] S743: By updating W gC and W sC , the gating mechanism can adaptively fuse the syntactic and semantic feature vectors of the code, improving the accuracy of calculating the similarity score.

[0103] Compared with the prior art:

[0104] The present invention evaluates the overall similarity of the code by calculating the similarity between the code syntax and semantic features; in terms of extracting the code syntax feature vector, when aggregating features for each node in the syntax tree, the depth of the node and the features of its sibling nodes are considered, avoiding feature loss during the aggregation of the tree structure by conventional methods; in terms of extracting the code semantic feature vector, by introducing an adaptive forgetting factor, the most critical semantic information of the code can be retained during the extraction of the semantic feature vector, and the operation efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 is a schematic diagram of the process of the present invention;

[0106] Figure 2 is a schematic diagram of the Python code of the present invention;

[0107] Figure 3 is a schematic diagram of the syntax tree T c of the present invention;

[0108] Figure 4 is a schematic diagram of the adjacency matrix R C (left) and the adjacency matrix (right) of the present invention;

[0109] Figure 5 is an architecture diagram of the AST-GCN model of the present invention;

[0110] Figure 6 is a schematic diagram of the sibling matrix of the present invention;

[0111] Figure 7 is a schematic diagram of the depth matrix d C (left) and the depth weight matrix d' C (right) of the present invention;

[0112] Figure 8 is a schematic diagram of the multi-scale semantic feature adaptive fusion model of the present invention;

[0113] Figure 9 is a structural diagram of the F-BiSRU module of the present invention;

[0114] Figure 10 is a structural diagram of the FSRU unit of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0115] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0116] The present invention provides a code similarity detection method based on deep learning. Please refer to Figures 1-10 , which includes the following specific steps:

[0117] S1: Collect codes, preprocess the codes. The operations include deleting corresponding comment content, normalizing variable names and function names, unifying the code format, making any pairwise combinations of the processed codes, and rating the code similarity to obtain a triple <code A, code B, similarity y>. The set of all triples constitutes a code data set. Specifically, since the same processing needs to be performed on code A and code B in each triple, unless otherwise specified, the following steps are applicable to both code A and code B. For the convenience of expression, in the following steps, code C is used to uniformly represent code A and code B, that is, either code A or code B is set as code C;

[0118] S2: Use a code extraction tool to parse code C and generate a syntax tree, traverse the syntax tree to obtain the adjacency matrix and the depth weight matrix d' C , and then convert code C into a feature vector of a specified dimension through Word2Vec to obtain a feature matrix Z C ;

[0119] S2 includes but is not limited to the following embodiments:

[0120] Adopt a code extraction tool such as Joern or Antlr to extract the syntax tree of the code. These tools will first perform lexical analysis, decompose the code into basic word units, and form a lexical unit stream according to the order of words in the code. Then, start to perform syntax analysis on the lexical unit stream and gradually construct a syntax tree; the syntax tree is presented in a multi-way tree structure, and each node of the tree corresponds to different syntax elements;

[0121] There is a Python code for calculating the product as Figure 2 shown;

[0122] Parse the above code and construct a syntax tree T c , and adopt a pre-order traversal method to traverse T c , and during the traversal, number each node starting from 1. The obtained syntax tree T c is as Figure 3 shown;

[0123] For Figure 3 the T inc Perform a pre-order traversal. When visiting a node and its children, record the edges between the node and its children. The specific approach is as follows: form a pair of the parent node number and the child node number to represent the above-mentioned edge, and then add the pair to an edge list. For example, when traversing to the FunctionDeclarction node and the Param:[x,y] node, record the pair <2,3> in the edge list. Since the syntax tree is an undirected graph, <3,2> also needs to be recorded in the edge list.

[0124] Use Python libraries such as numpy to directly generate the adjacency matrix R from the edge list. C , since each node also needs to retain its own features when aggregating neighbor features according to the adjacency matrix, there should also be an edge connecting each node to itself, that is, in R C change the position corresponding to each node to itself to 1 to obtain the adjacency matrix As Figure 4 shown;

[0125] The data formats of the points and edges that the graph convolutional neural network can process must both be vector formats. Word2Vec can convert the points and edges in the syntax tree into vector representations. For example, for the FunctionDeclaration node, that is, node 1, Word2Vec will output a 32-dimensional sequence [1.2, 4.0, …, 1.5] representing the features of node 1. Arrange the feature sequences of all nodes in the order of node numbers to form a matrix. Each row of the matrix corresponds to the feature sequence of a node, and the number of columns is the dimension of the feature vector. Obtain the feature matrix Z C :

[0126]

[0127] where i ∈ (1, 2, …, 15) represents the feature sequence of the node with node number i;

[0128] Among them, Word2Vec is a method in the Python toolkit Gensim. Its purpose is to learn the distributed representation of words through a neural network model and map words to a low-dimensional vector space. In this vector space, semantically similar words are close in distance, so as to better capture the semantic and syntactic relationships between words.

[0129] Note: numpy (Numerical Python) is a basic library for scientific computing in the Python language, providing Python with efficient and flexible multi-dimensional array processing capabilities and simplifying the data processing process. The array() method provided by numpy can convert the information in the edge list into an adjacency matrix.

[0130] S3: Feed and Z C into the AST-GCN model, which aggregates the syntactic features of the syntax tree of code C to obtain a syntactic feature matrix Finally, through the multi-head self-attention layer, is feature-enhanced to obtain the syntactic feature vector V of code C gC ; The structure of the AST-GCN model is as Figure 5 shown;

[0131] Among them, the specific steps of S3 are as follows:

[0132] S31: Construct the AST-GCN model for extracting the syntactic feature matrix of code C Its AST-GCN model consists of an input layer, a feature update layer, and an output layer; The AST-GCN model fuses the features of each node itself in the syntax tree with the features of its neighbor nodes and sibling nodes to obtain a feature sequence containing the local syntactic features of the code Finally, through a linear transformation, the syntactic feature matrix is obtained

[0133] Among them, the Graph Convolutional Network (GCN) is a deep learning method specifically used to process graph-structured data; The core idea of GCN is to define a convolution operation on the graph structure, so as to be able to learn the feature representation of the nodes in the graph; It extends the convolution operation from traditional grid data (such as the pixel grid of an image, the sequence of text) to an irregular graph structure, enabling the use of the neighborhood information of nodes to update the features of nodes; GCN often consists of multiple convolutional layers. When updating the features of a certain node, each layer will aggregate the features of this node and output the updated feature sequence as the input of the next layer;

[0134] Among them, the specific steps of S31 are as follows:

[0135] S311: Traverse the syntax tree of code C hierarchically to obtain a sibling matrix and a depth weight matrix d' C ;

[0136] Among them, the specific steps of S311 are as follows:

[0137] S3111: Traverse the syntax tree hierarchically. For any node r encountered during the traversal, determine the sibling nodes of r according to the adjacency relationship between nodes, and thus obtain a sibling matrix In Stores all sibling relationships between syntax tree nodes, that is, if node p and node q are siblings, then there is x p,q 、 x p,q =x q,p =1; otherwise x p,q =x q,p =0;

[0138] Among them, for Figure 3 T in c perform a level traversal starting from node 1, using a queue to assist in the traversal; add node 1 to the queue; traverse all its son nodes; for these son nodes, which are siblings of each other, mark these nodes as siblings. For example, if node 2 and node 11 are siblings, then there is x 2,11 、 x 2,11 =x 11,2 =1, and at the same time add nodes 2 and 11 to the queue for subsequent traversal of their son nodes to obtain the sibling matrix As Figure 6 shown;

[0139] S3112: When traversing the syntax tree, record the depth of each node in the syntax tree in the depth matrix d C and adjust the weight coefficient during feature aggregation according to the preset maximum critical depth m. For any node i in d C whose depth is greater than m, the value of i is replaced by where α is a scaling coefficient and d i represents the depth of node i; while for any node j whose depth is less than or equal to m, the value of j is replaced by 1. After completing the above adjustment of the weight coefficient, finally obtain the depth weight matrix d' C ;

[0140] Among them, for Figure 3 T in c perform a level traversal, and record the depth of each node in the syntax tree with the depth matrix d C ; d C is a diagonal matrix, and the values on the diagonal represent the depths of the corresponding nodes; in this embodiment, set the critical maximum depth m to 2, then for the values in d C that are greater than 2, replace them with where α is set to 1; for the values less than or equal to 2, replace them with 1; obtain the depth weight matrix d' C , as Figure 7 shown;

[0141] S312: Take '

[0142] dC , and Z C An input feature update layer, which is composed of L convolutional layers. Suppose the feature update is performed on the (l + 1)-th layer (l ∈ (0, 1, …, L - 1)) of the syntax tree in the feature update layer, then the feature matrix output by this layer is composed of the addition of three parts of features, namely the sibling feature matrix the hierarchical neighbor feature matrix and the feature matrix output by the previous layer of the feature update layer

[0143] Among them, the specific steps of S312 are as follows:

[0144] S3121: Calculate the sibling feature matrix Using and to perform a feature aggregation operation to obtain

[0145]

[0146] Among them is the inverse matrix of the degree matrix, and W s is the weight matrix;

[0147] S3122: For the nodes deeper in the syntax tree, in order to reduce the influence of the features of these nodes on the overall similarity of the code, multiply on the left by the depth weight matrix d' C and then perform feature aggregation to obtain the hierarchical neighbor feature matrix

[0148]

[0149] Among them, W n is the weight matrix;

[0150] S313: Add and to obtain the updated feature matrix on the (l + 1)-th layer of the feature update layer

[0151]

[0152] where ReLU is the activation function;

[0153] S313 includes but is not limited to the following embodiments:

[0154] Figure 3 T c The feature update process of the first layer of the feature update layer:

[0155] 1. Calculate the sibling feature matrix

[0156]

[0157] 2. Calculate the hierarchical neighbor feature matrix

[0158] Let the maximum critical depth m be 2, and the scaling factor α be set to 1 to obtain the depth weight matrix d' C , calculate

[0159]

[0160] 3. Calculate the feature sequence output by the first layer of the feature update layer

[0161]

[0162] 4. At this point, Z C After passing through the first layer of the feature update layer, it is updated to Take As the input of the second layer to continue the feature update, repeat the above operations until the final layer;

[0163] S314: After calculating through the final layer of the feature update layer, output the feature matrix

[0164]

[0165] Among them i ∈ (1, 2, …, n) is the updated feature sequence output by each node feature sequence C in Z After the final layer of the feature update layer, perform a linear transformation on to obtain the syntactic feature matrix

[0166] S32: Use the multi-head self-attention mechanism to calculate the attention weights for , concatenate the syntactic feature sequences output by each attention head, and then obtain the syntactic feature vector V through a linear transformation gC ;

[0167] Among them, the self-attention mechanism is a deep learning mechanism. For an input sequence, the input sequence is transformed into query vectors, key vectors, and value vectors through a learnable weight matrix; then, the dot product of the query vector and the key vector is calculated and scaled to obtain attention scores, which are normalized by Softmax to obtain attention weights, and then the value vectors corresponding to the attention weights are weighted and summed to obtain the output; in this way, the dependencies between elements in the sequence can be captured and are not restricted by distance; the multi-head self-attention mechanism repeats the above self-attention process multiple times, each time using a different weight matrix, and each self-attention head will output a sequence, and then these sequences are concatenated and then mapped back to the original input dimension through a linear transformation;

[0168] A linear transformation performs a dimensional transformation on vectors in a vector space. This kind of transformation does not change the linear relationship between vectors and generally maintains the basic structure and properties of the vector space; usually, a linear transformation can be achieved by right-multiplying a matrix of a specific dimension to match the dimension of the next model; unless otherwise specified in the following steps, the purpose of the linear transformation operation is to match the vector dimension with the input dimension of the model;

[0169] S4: Perform word segmentation on code C, and use a word segmentation tool to split code C into several lexical units to obtain a sequence of lexical units X C ;

[0170] Among them, the tool generally used for word segmentation of code is Tokenizer. Tokenizer can split the code into multiple lexical units (called tokens); it usually operates based on predefined rules or learned patterns;

[0171] S5: Construct a multi-scale semantic feature adaptive fusion model MSF (Multi Semantic Fusion) for multi-scale fusion and extraction of code semantic features. After inputting X C into the MSF model, a feature vector V representing the code semantics is obtained sC ; The architecture of the multi-scale semantic adaptive fusion model is as Figure 8 shown;

[0172] Among them, the specific steps of S5 are as follows:

[0173] S51: First, input X C into CodeBERT, perform a concat operation on the semantic sequences output by each self-attention head of CodeBERT to obtain a feature vector that preliminarily represents the code semantics, and then obtain the semantic vector v of code C through a linear transformation C ;

[0174] Among them, CodeBERT is a pre-trained model based on the Transformer encoder architecture. This model contains multiple self-attention heads, which are specifically used to process program code and natural language. Each self-attention head outputs a semantic sequence. Therefore, a concat operation is required to splice these sequences to obtain a complete semantic vector;

[0175] S52: Downsample v C to obtain a semantically coarse-grained sequence Then, input v C and into the bidirectional simple recurrent unit with an adaptive forgetting factor, F-BiRSU (Adapted Forgetable BiSRU), to obtain the semantic feature sequence H of code C C ;

[0176] Among them, the specific steps of S52 are as follows:

[0177] S521: Use max pooling to downsample v C to obtain a semantically coarse-grained sequence For any 1 ≤ i ≤ m, its calculation formula is as follows:

[0178]

[0179] Among them, is the semantic vector of code C, k is the size of the max pooling window, and n is an integer multiple of k;

[0180] S521 includes but is not limited to the following embodiments:

[0181] Assume that the semantic vector x = [3, 5, 9, 11, 13, 17] and the downsampling window k = 3. Then:

[0182] When i = 1, (1 - 1)k + 1 = 1, 1 * k = 3, then

[0183] When i = 2, (2 - 1)k + 1 = 4, 2 * k = 6, then

[0184] When i = 3, (3 - 1)k + 1 = 7, 3 * k = 9, but the original sequence has only 8 elements, so the last element is taken here; then

[0185] In summary, the downsampled coarse-grained sequence x d= [7, 13, 17, 0, 0, 0];

[0186] S522: Construct an F-BiSRU module for extracting the semantic feature sequence H from v C and The overall structure of the F-BiSRU module is as shown in C , and the F-BiSRU module is composed of a forward FSRU network and a backward FSRU network. Each network contains several FSRU units, and the structure of the FSRU unit is as shown in Figure 9 ; Figure 7 ;

[0187] The principle of the F-BiSRU module is as follows: When taking v as the input of the F-BiSRU module, it is necessary to divide v C into several time series sequences respectively, that is, to obtain v C . The FSRU unit in the F-BiSRU module is used to process a pair of time series sequences at any time t and the processing process is the same: Scale and perform feature fusion on and and to obtain a multi-scale information fusion vector , and introduce the output of the hidden state at the previous moment and to be concatenated as the parameter for forgetting gate update, and finally obtain the output of the hidden state at the current moment

[0188] Among them, the designed F-BiSRU module of the present invention is improved on the basis of the existing BiSRU model. The BiSRU model is briefly introduced below;

[0189] BiSRU (Bidirectional Simple Recurrent Unit), that is, a bidirectional simple recurrent unit, is a neural network structure developed on the basis of SRU (Simple Recurrent Unit); it is composed of a forward SRU network and a backward SRU network, and each SRU network is composed of several SRU units; BiSRU combines the idea of bidirectional processing of sequence data and can consider both the forward and backward information of the sequence at the same time, so as to model the sequence data more comprehensively; in the forward SRU network, it updates the values of its forgetting gate f, reset gate r, and memory cell c by combining the forward input sequence, weight parameters, and bias values, and then combines the forgetting gate f, reset gate r, candidate memory cell c, and the tanh function to calculate the value of the hidden layer vector h. The calculation process of the gating unit and memory cell unit in the backward SRU network is the same as that in the forward SRU network;

[0190] Among them, the specific steps of S522 are as follows:

[0191] S5221: First, define the gating vector I C , which is used to control and v C the fusion ratio of the time series input at each moment. The calculation formula of I C is as follows:

[0192]

[0193] where Sigmoid is the activation function, W t is the weight matrix, is the bias vector, represents concatenating v C and the time series at time t and , and then performing multi-scale information fusion on and according to I C to obtain

[0194]

[0195] where ⊙ is the element-wise multiplication operation;

[0196] S5221 includes but is not limited to the following embodiments:

[0197] Element-wise multiplication refers to multiplying the elements in the corresponding positions of two matrices or vectors with the same dimensions; for example, there are two vectors a = [a1, a2, a3] and b = [b1, b2, b3], and the result of their element-wise multiplication is a⊙b = [a1×b1, a2×b2, a3×b3]; if it is a matrix, it is also multiplying the elements in the corresponding positions one by one;

[0198] S5222: When calculating the output of the forgetting gate at time t of the forward FSRU network, introduce the hidden state at the previous moment Concatenate with as the parameter for updating . The specific calculation formula is as follows:

[0199]

[0200] where W f is the weight matrix, b f is the bias vector, is the hidden state of the forward FSRU network at the previous moment, is a statistical characteristic function based on the variance, where θ is a coefficient used to scale the value;

[0201] The reset gate r t , memory cell c t and the forward hidden state output at time t of the forward FSRU network are calculated as follows:

[0202]

[0203]

[0204]

[0205] The backward hidden state output at time t of the backward FSRU network is calculated in the same way as the forward hidden state output at time t of the forward FSRU network, but when calculating the forget gate , the previous hidden state

[0206]

[0207] S523: Concatenate and through the concat operation to obtain and then obtain the semantic feature sequence H C , that is:

[0208]

[0209] where

[0210] S53: Use the self-attention layer to further enhance the semantics of H C so that H C can incorporate information from other positions in the entire sequence and better capture the global semantic information of the code, and finally obtain the semantic feature vector V sC of code C;

[0211] S6: Use the gating mechanism to adaptively fuse the syntax feature vector V gC and the semantic feature vector V sC of code C to obtain the fusion vector V fusC ;

[0212] Among them, the specific steps of S6 are as follows:

[0213] S61: Design a linear layer composed of a learnable syntax matrix W gC and a bias value b gC to generate the syntax gating vector g of code C gC :

[0214] g gC = ReLU(W gC V gC + b gC )

[0215] Similarly, design a linear layer composed of a learnable semantic matrix W sC and a bias value b sC to generate the semantic gating vector g of code C sC :

[0216] g sC = ReLU(W sC V sC + b sC )

[0217] Calculate the fusion weight α of the syntax feature vector of code C C and the fusion weight β of the semantic feature vector C :

[0218]

[0219] β C = 1 - α C

[0220] S62: Use α C , β C to perform weighted summation on V gC , V sC to obtain the final fusion vector V of code C fusC :

[0221] V fusC = α C V gC + β C V sC ;

[0222] Among them, since code A or code B is set as code C in S1, so V fusC actually represents the fusion vector V of code A fusA , the fusion vector V of code B fusB , that is, after S62, V fusA and V fusB have been obtained;

[0223] S7: Concatenate V fusA and V fusB to obtain the vector V to be measuredt , input V t into a three - layer perceptron for similarity comparison, and output a similarity score Calculate the loss value loss between the output and the true value, and update the learnable grammar matrix W in the reverse direction gC and the learnable semantic matrix W sC ;

[0224] Among them, backpropagation in deep learning is achieved through the chain rule of differentiation; after constructing a neural network model, compare the difference between the output value of the model and the true value to obtain the loss value, and then start from the output layer, and update the parameters of each layer of the model (such as weights and biases) in the reverse direction according to the chain rule of differentiation; this method uses the loss value to update the learnable grammar matrix and the learnable semantic matrix in the gating mechanism in the reverse direction to reduce the loss value between the finally output similarity score and the true value;

[0225] Among them, the specific steps of S7 are as follows:

[0226] S71: Concatenate the fusion vectors V fusA , V fusB to obtain the vector V to be measured t :

[0227]

[0228] where represents the concatenation operation;

[0229] S72: Input V t into a three - layer perceptron for training, and output a value between 0 and 1 as the similarity score between code A and code B

[0230] Z1 = ReLU(W1V t + b1)

[0231] Z2 = ReLU(W2Z1 + b2)

[0232] Z3 = ReLU(W3Z2 + b3)

[0233]

[0234] where W1, W2 and W3 are weight matrices, b1, b2 and b3 are bias vectors, and Z1, Z2 and Z3 are the feature representations output by the first, second and third layers in the three - layer perceptron respectively;

[0235] S72 includes but is not limited to the following embodiments:

[0236] Suppose that in the first training task, the fusion vector V of code A and code B is calculated fusA and V fusB , and then the similarity score between code A and code B is calculated

[0237] (1) Calculate the fusion vector V of code A fusA :

[0238] Let the syntax feature vector of code A semantic feature vector

[0239] Calculate the syntax gating vector g gA :

[0240] Let the learnable syntax matrix bias value Calculate g gA :

[0241]

[0242] Calculate the semantic gating vector g sA :

[0243] Let the learnable semantic matrix bias value Calculate g sA :

[0244]

[0245] Calculate the fusion weights α A and β A :

[0246]

[0247] β A = 1 - α A ≈ [0.57, 0.5]

[0248] Calculate the fusion vector V fusA :

[0249]

[0250] (2) Calculate the fusion vector V of code B fusB :

[0251] Using the same steps as in (1), the fusion vector of code B is calculated

[0252] (3) Calculate the similarity score between code A and code B

[0253] Let VfusA and V fusB Concatenate them to obtain the vector V to be measured t :

[0254]

[0255] Calculate the similarity score

[0256] Input V t into a three-layer perceptron to obtain the similarity score between the code A and code B of this training task

[0257] S73: Calculate the loss value loss between the similarity score and the true similarity value. The calculation formula is as follows:

[0258]

[0259] where y i is the true similarity value between the code A and code B of this training. The closer its value is to 1, the more similar code A and code B are;

[0260] S74: Use the loss value loss to perform backpropagation updates on the learnable syntax matrix W gC and the learnable semantic matrix W sC ;

[0261] The specific steps of S74 are as follows:

[0262] S741: The update process of W gC is as follows:

[0263]

[0264] S742: The update process of W sC is as follows:

[0265]

[0266] where θ represents the partial derivative, and η1, η2 are the learning rates;

[0267] S743: By updating W gC and W sC , the gating mechanism can adaptively fuse the syntax and semantic feature vectors of the code, improving the accuracy of calculating the similarity score;

[0268] Among them, in S74, since it has been declared in S1 that code C is used to represent the operation process of code A and code B, so W gC substantially represents the learnable syntax matrices W gA of code A and code B, WgB ; W sC represents the learnable semantic matrices W of code A and code B sA , W sB ;

[0269] S741: For the learnable syntax matrices W of code A and code B gA and W gB The update process is as follows:

[0270]

[0271]

[0272] S742: For the learnable semantic matrices W sA and W sB The update process is as follows:

[0273]

[0274]

[0275] where represents the partial derivative, and η1, η2 are the learning rates;

[0276] S743: By updating W gA , W gB as well as W sA , W sB , the gating mechanism can adaptively fuse the syntax and semantic feature vectors of the code, improving the accuracy of calculating the similarity score;

[0277] Among them, the Three-Layer Perceptron is an artificial neural network, which consists of an input layer, a hidden layer, and an output layer. Each layer outputs a feature representation after linear transformation and activation by an activation function, which is used as the input for the next layer; the hidden layer can increase the complexity and expressiveness of the model, enabling it to handle more complex non-linear problems. This method uses the Three-Layer Perceptron to calculate the vector similarity;

[0278] S8: Apply the trained code similarity detection model to the code similarity detection task, calculate the similarity of any actual input code pair <code X, code Y>, and obtain the similarity score of the code pair.

[0279] In summary, the present invention proposes an AST-GCN model adapted to the characteristics of the tree structure. When traversing the syntax tree of the code and updating the node features, this model will fuse the features of the sibling nodes of each node and adopt different aggregation weights according to the depth of the syntax tree. Finally, the syntax feature vector of the code is obtained by inputting into the multi-head self-attention layer.

[0280] The present invention proposes a multi-scale semantic adaptive fusion model MSF. This model initially extracts the semantic vectors of code using CodeBERT, then downsamples these semantic vectors through a max pooling layer, and further performs information fusion and feature extraction on them through the F-BiSRU module. Finally, the semantic feature vectors integrating multi-scale information of the code are calculated by inputting them into the self-attention layer.

[0281] The present invention assigns learnable weight matrices to the syntax and semantic feature vectors of two pieces of code respectively, enabling the values of the learnable weight matrices to be updated backward according to the final loss value, so as to adaptively fuse the syntax and semantic feature vectors of the code.

[0282] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and its components can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the features in the embodiments disclosed in the present invention can be combined with each other in any way. The exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A code similarity detection method based on deep learning, characterized in that The specific steps are as follows: S1: Collect codes and preprocess them. The operations include deleting corresponding comment content, normalizing variable names and function names, and unifying the code format. Combine the processed codes in any pairwise manner and perform code similarity rating to obtain a triple <Code A, Code B, similarity y>. The set of all triples constitutes a code dataset, where Code A or Code B is set as Code C; S2: Parse the code C using a code extraction tool to generate a syntax tree, traverse the syntax tree to obtain the adjacency matrix of the code C and the depth weight matrix d' C , and then convert the code C into a feature vector of a specified dimension through Word2Vec to obtain the feature matrix Z C ; S3: Input and Z C into the AST-GCN model, and the AST-GCN aggregates the syntactic features of the syntax tree of code C to obtain a syntactic feature matrix Finally, the multi-head self-attention layer is used to perform feature enhancement to obtain the syntactic feature vector V of code C gC ; S4: Perform word segmentation on code C, and use a word segmentation tool to split code C into several lexical units, obtaining a lexical unit sequence X C ; S5: Construct a multi-scale semantic feature adaptive fusion model MSF for multi-scale fusion and extraction of code semantic features, and input X C into the MSF model to obtain a feature vector V representing the code semantics sC ; S6: Use a gating mechanism to adaptively fuse the C syntax feature vector V gC of the code and the semantic feature vector V sC to obtain a fused vector V fusC ; S7: Concatenate V fusA and V fusB to obtain the vector V t to be measured. Input V t into a three-layer perceptron for similarity comparison and output a similarity score Calculate the loss value loss between the calculated value and the true value, and update the learnable grammar matrix W gC and the learnable semantic matrix W sC in the reverse direction; S8: Apply the trained code similarity detection model to the code similarity detection task, calculate the similarity of any actually input code pair <Code X, Code Y>, and obtain the similarity score of the code pair; The specific steps of S3 are as follows: S31: Construct an AST-GCN model for extracting the syntax feature matrix of code C The AST-GCN model consists of an input layer, a feature update layer, and an output layer; the AST-GCN model fuses the features of each node itself in the syntax tree with the features of its neighbor nodes and sibling nodes to obtain a feature sequence containing the local syntax features of the code Finally, a syntax feature matrix is obtained through a linear transformation S32: Use the multi-head self-attention mechanism to calculate the attention weights, concatenate the syntactic feature sequences output by each attention head, and then obtain the syntactic feature vector V through a linear transformation gC .

2. The method for detecting code similarity based on deep learning according to claim 1, characterized in that The specific steps of S31 are as follows: S311: Traverse the syntax tree of code C in hierarchical order to obtain the sibling matrix and the depth weight matrix d' C ; S312: Take d′ C , and Z C as inputs to the feature update layer, which consists of L convolutional layers. Assume that the syntax tree performs feature update at the (l + 1)-th layer (l ∈ (0, 1, …, L - 1)) in the feature update layer. Then the feature matrix output at this layer is composed of the sum of three parts of features, namely the sibling feature matrix the hierarchical neighbor feature matrix and the feature matrix output at the previous layer of the feature update layer S313: Add and to obtain the updated feature matrix at the (l + 1)-th layer of the feature update layer where ReLU is the activation function; S314: After the calculation of the final layer of the feature update layer, an output feature matrix is output. where i ∈ (1, 2, …, n) is for Z C each node feature sequence in At the final layer output of the feature update layer, the updated feature sequence is subjected to a linear transformation to obtain a syntax feature matrix 3. The method for detecting code similarity based on deep learning according to claim 2, wherein The specific steps of S311 are as follows: S3111: Traverse the syntax tree hierarchically. For any node r encountered during the traversal, determine the sibling nodes of r based on the adjacency relationship between nodes, thereby obtaining the sibling matrix Store all sibling relationships between syntax tree nodes in , that is, if nodes p and q are sibling nodes, then there is x p,q , x p,q = x q,p = 1; otherwise x p,q = x q,p = 0; S3112: When traversing the syntax tree, record the depth of each node in the syntax tree in the depth matrix d C where d C is a diagonal matrix, and the values on its diagonal are the depths of each node; adjust the weight coefficient during feature aggregation according to the preset maximum critical depth m. For any node i in d C whose depth is greater than m, the value of i is replaced with where α is the scaling coefficient, and d i represents the depth of node i; for any node j whose depth is less than or equal to m, the value of j is replaced with 1. After completing the above adjustment of the weight coefficient, the depth weight matrix d' C is finally obtained; The specific steps of S312 are as follows: S3121: Calculate the sibling feature matrix Use and to perform a feature aggregation operation to obtain where is the inverse matrix of the degree matrix, and W s is the weight matrix; S3122: For nodes deeper in the syntax tree, in order to reduce the impact of the features of these nodes on the overall code similarity, left-multiply by the depth weight matrix d' and then perform feature aggregation to obtain the hierarchical neighbor feature matrix C ​ Among which W n is the weight matrix.

4. A code similarity detection method based on deep learning according to claim 1, characterized in that The specific steps of S5 are as follows: S51: First, input X C into CodeBERT, perform a concat operation on the semantic sequences output by each self-attention head of CodeBERT to obtain a feature vector that preliminarily represents the code semantics, and then obtain the semantic vector v of code C through a linear transformation C ; S52: Downsample v C to obtain a semantically coarser-grained sequence Then input v C and into the bidirectional simple recurrent unit with an adaptive forgetting factor, F-BiRSU, to obtain the semantic feature sequence H of code C C ; S53: Further semantically enhance H using a self-attention layer C so that H C can incorporate information from other positions in the entire sequence, better capture the global semantic information of the code, and finally obtain the semantic feature vector V of code C sC .

5. A method for detecting code similarity based on deep learning according to claim 4, characterized in that, The specific steps of S52 are as follows: S521: Perform downsampling on v using max pooling to obtain a semantically coarse-grained sequence C For any 1 ≤ i ≤ m, its calculation formula is as follows: ​ Among them, is the semantic vector of code C, k is the size of the maximum pooling window, and n is an integer multiple of k; S522: Construct an F-BiSRU module for extracting the semantic feature sequence H from v C and C , where the F-BiSRU module consists of a forward FSRU network and a backward FSRU network, and each network contains a number of FSRU units;​ S523: Concatenate and using the concatenation operation concat to obtain and then obtain the semantic feature sequence H through a linear transformation C , that is: Among them The specific steps of S522 are as follows: S5221: First, define the gating vector I C , which is used to control and v C 's fusion ratio of the input time series at each moment. The calculation formula of I C is as follows: where Sigmoid is the activation function, W t is the weight matrix, is the bias vector, represents concatenating v C and at the time series at time t and then performing splicing, and then and performing multi-scale information fusion according to I C to obtain where ⊙ is the element-wise multiplication operation; S5222: When calculating the output of the forgetting gate at time t of the forward FSRU network, introduce the hidden state of the previous moment and splice it with as the parameter for updating . The specific calculation formula is as follows: The splicing is used as the parameter for updating : Among which W f is the weight matrix, b f is the bias vector, is the hidden state at the previous moment of the forward FSRU network, is based on the statistical characteristic function of the variance of, θ is the coefficient used to scale the magnitude of the value; Reset gate r of the forward FSRU network at time t t , memory cell c t and the forward hidden state output at time t The calculation process is as follows: The reverse hidden state output by the reverse FSRU network at time t The calculation steps are the same as those for calculating the forward hidden state output by the forward FSRU network at time t, but when calculating the forget gate the hidden state of the previous moment of the reverse FSRU network is introduced 6. The method for detecting code similarity based on deep learning according to claim 1, characterized in that, The specific steps of S6 are as follows: S61: Design a linear layer composed of a learnable grammar matrix W gC and a bias value b gC to generate a grammar gating vector g of the code C gC : g gC = ReLU(W gC V gC + b gC ) Similarly, design a linear layer composed of a learnable semantic matrix W sC and a bias value b sC to generate the semantic gating vector g of the code C sC : g sC = ReLU(W sC V sC + b sC ) Calculate the fusion weight α of the syntactic feature vector of code C C and the fusion weight β of the semantic feature vector C : β C =1-α C S62: Using α C , β C to perform weighted summation on V gC , V sC to obtain the final fusion vector V of code C fusC : V fusC = α C V gC + β C V sC ; Among them, since code A or code B is set as code C in S1, V fusC actually represents the fusion vector V fusA of code A and the fusion vector V fusB of code B at the same time, that is, after S62, V fusA and V fusB have been obtained.

7. A method for detecting code similarity based on deep learning according to claim 6, characterized in that, The specific steps of S7 are as follows: S71: Concatenate the fusion vectors V fusA and V fusB to obtain the vector V t to be measured: Among them represents a splicing operation; S72: Input V t into a three-layer perceptron for training, and output a value between 0 and 1 as the similarity score between Code A and Code B Z1 = ReLU(W1V t + b1) Z2 = ReLU(W2Z1 + b2) Z3 = ReLU(W3Z2 + b3) where W1, W2, and W3 are weight matrices, b1, b2, and b3 are bias vectors, and Z1, Z2, and Z3 are the feature representations output by the first, second, and third layers in the three-layer perceptron respectively; S73: Calculate Calculate the loss value loss between the predicted similarity value and the true similarity value. The calculation formula is as follows: where y i is the true similarity value between code A and code B in this training. The closer its value is to 1, the more similar code A and code B are; S74: Use the loss value loss to perform backward updates on the learnable syntax matrix W gC and the learnable semantic matrix W sC for backward updates 8. A method for detecting code similarity based on deep learning according to claim 7, characterized in that The specific steps of S74 are as follows: S741: Update process for W gC is as follows: S742: Update process for W sC is as follows: where denotes the partial derivative, and η1, η2 are the learning rates; S743: By updating W gC and W sC , the gating mechanism can adaptively fuse the syntactic and semantic feature vectors of the code, improving the accuracy of calculating the similarity score.

Citation Information

Patent Citations

  • Smart contract code clone detection method based on AST multi-dimensional feature fusion

    CN115422541A

  • Source code vulnerability static detection and positioning method based on graph neural network

    CN115935367A

  • Software defect prediction method based on abstract syntax tree

    CN115982037A

  • Deep learning-driven semantic grammar interaction code annotation generation method and device

    CN117891460A

  • Intelligent contract multiplexing hierarchical detection method based on grammar and semantic separation

    CN118626376A

Cited By

  • Intelligent comparison and analysis method and system for similarity of examination answer codes

    CN120631735A

  • FPGA code similarity detection method and device based on weighted directed graph

    CN120670323A