Multi-view comparative learning cross-language source code representation method

Through multi-perspective comparative learning method, combined with meta-learning and graph neural network, a cross-language source code representation model is built, which solves the problems of ignoring language-specific information and excessive smoothing in graph representation in the existing technology, and realizes more effective cross-language source code representation.

CN120144100AActive Publication Date: 2025-06-13BEIJING INST OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510168743.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-13
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

Existing cross-language source code representation methods mainly focus on learning a unified model of shared parameters between different languages, ignoring the model of language-specific information, and graph neural networks are difficult to capture multi-hop neighbors, resulting in excessive smoothing problems in the model extracting global structural information.

Method used

Multi-view comparison learning method is adopted, Transformer's learnable parameters are initialized through meta-learning, code feature heterogeneous graphs are constructed and GCN is aggregated, and graph node embedding is generated by combining meta-paths and hierarchical attention mechanisms. Finally, loss function training is established through comparison learning to generate a cross-language source code representation model.

Benefits of technology

Effectively integrate language-specific information into the Transformer structure, improve the model's ability to represent cross-language source code, avoid excessive smoothing in graph representation learning, and improve feature extraction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144100A_ABST
    Figure CN120144100A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-view comparative learning cross-language source code representation method, and belongs to the technical field of computer software information. The method comprises the following steps: firstly, initializing Transform learnable parameters according to a programming language type by utilizing a meta-learning method; secondly, constructing a code feature heterogeneous graph by depending on information such as grammar and structure of a source code fragment, and aggregating the heterogeneous graph by using GCN to obtain graph node embedding; then aggregating node same-hop neighborhood information according to a meta path to obtain serialized representation of graph nodes, and generating graph node embedding by using Transform and a hierarchical attention mechanism; and finally, establishing a contrast loss function according to node embedding under the two perspectives, and training to generate a cross-language source code representation model. Aiming at the problems that the representation effect is influenced by application of universal code features of a programming language and the model is over-smooth in the existing method, the specific information of the programming language is extracted, and the model is built by utilizing multi-view graph node embedding, so that the cross-language code representation effect is improved, and the accuracy of code abstract, source code vulnerability detection and the like is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-view contrast learning cross-language source code representation method, belonging to the field of computer software information technology. Background Art

[0002] Source code representation learning converts source code into a form that can be understood and processed by machine learning models, such as feature vectors like word embeddings and graph embeddings, and is widely used in many software engineering tasks such as code summarization and code completion. With the continuous increase in software scale and complexity, and the gradual popularization of multi-language integration platforms, cross-language code conversion and reuse can greatly improve the work efficiency of developers. How to effectively construct a unified cross-language code representation method has important theoretical significance and practical value. The current mainstream cross-language source code representation methods can be divided into tree-based representation methods and graph-based representation methods.

[0003] 1. Tree-based Representation Method

[0004] The tree-based representation method represents the source code in the form of an Abstract Syntax Tree (AST), and then uses tree-based neural network models (such as ASTNN, Tree-LSTM, etc.) to learn vector representations by recursively calculating node embeddings from bottom to top. Each node in the AST represents a syntax unit in the source code, such as a statement, an expression, a variable, or a function declaration, etc., and the edges represent the association relationships between these syntax units, such as parent-child relationships, sibling relationships, etc. In the cross-language code representation task, existing methods construct code representations based on the AST, and unify the keyword representations in the ASTs of different programming languages through means such as vocabulary mapping. Existing cross-language code representation methods mainly perform representation learning on language-independent information such as the control structure and processing logic of the code, and focus on constructing a unified cross-language code representation method that learns shared parameters for different languages, and pay less attention to the impact of language-specific information on the model's learning of code representations. The model is difficult to learn the syntax rules and character representations of code in a specific programming language, which limits the feature extraction effect of the deep learning encoder.

[0005] 2. Graph-based Representation Method

[0006] The graph representation-based method first constructs a composite graph structure that integrates multiple code representation methods, such as abstract syntax trees, control flow graphs (CFGs), program dependence graphs (PDGs), etc. Then, it uses graph neural network models (such as GAT, GCN, GGNN, etc.) to extract graph features for code representation. The graph-based code representation method uses code statements, keywords, etc. as nodes and the control, dependence, etc. relationships between codes as edges. By adding or deleting nodes or edges in the graph, it realizes the integration of multiple code representations to adapt to different downstream tasks. Graph neural networks have a powerful local information aggregation ability due to the message passing mechanism. However, limited by the over-smoothing problem, they are difficult to effectively extract global information, which restricts the code representation ability of the model.

[0007] In summary, the existing cross-language source code representation methods mainly have the following problems: (1) Focusing on learning a unified model with shared parameters between different languages while ignoring the modeling of language-specific information; (2) The over-smoothing problem that graph neural networks are difficult to capture multi-hop neighbor points to extract global structural information. Therefore, the present invention proposes a multi-perspective contrastive learning cross-language source code representation method. Summary of the Invention

[0008] The object of the present invention is to propose a multi-perspective contrastive learning cross-language source code representation method for the problems that the existing cross-language source code representation methods are affected by the general code features of programming languages in terms of representation effect and the model has an over-smoothing problem.

[0009] The design principle of the present invention is as follows: First, use the meta-learning method to initialize the learnable parameters of the Transformer according to the programming language type of the input code; secondly, construct a code feature heterogeneous graph depending on the syntax, structure, etc. information of the source code fragment, and use GCN to aggregate the heterogeneous graph to obtain graph node embeddings; then, aggregate the node homophily neighborhood information according to the meta-path to obtain the serialized representation of the graph nodes, and use the Transformer and hierarchical attention mechanism to generate graph node embeddings; finally, establish a contrastive loss function based on the node embeddings from two perspectives to train and generate a cross-language source code representation model.

[0010] The technical solution of the present invention is realized through the following steps:

[0011] Step 1, use the meta-learning method to initialize the learnable parameters of the Transformer self-attention module according to the programming language type.

[0012] Step 2, construct a code feature heterogeneous graph for the source code fragment that includes code syntax, control, and dependence relationships, and use GCN to aggregate the heterogeneous graph to obtain graph node embeddings.

[0013] Step 2.1, generate AST, PDG, and CFG representations of source code in different languages using the open-source tools Joern, Soot, and ANTLR.

[0014] Step 2.2, expand based on the AST and incorporate the PDG and CFG. Traverse the edge sets of the PDG and CFG graphs, and add edges between the corresponding nodes in the AST according to the index information of the two end nodes, generating a heterogeneous graph with two types of structural information: control and dependency.

[0015] Step 2.3, perform node feature transformation to unify the feature dimensions of different types of nodes in the heterogeneous graph.

[0016] Step 2.4, use GCN to aggregate the heterogeneous graph to obtain node embeddings from the graph perspective.

[0017] Step 3, aggregate the co-hop neighborhood information of nodes according to the meta-path to obtain the serialized representation of graph nodes, and use Transformer and hierarchical attention mechanism to generate node embeddings from the serialized perspective.

[0018] Step 3.1, use the meta-path aware strategy to generate the representation vectors of graph nodes under each meta-path. First, obtain the neighbor nodes of different hop numbers of nodes under different meta-paths, then regard the neighbors from the same hop as a group, and perform information aggregation within the group, and splice to obtain the representation vector of the node under the meta-path.

[0019] Step 3.2, use Transformer to encode the node representation vectors to further explore the semantic interaction between different hop neighborhoods of nodes under the same meta-path.

[0020] Step 3.3, use the hierarchical attention mechanism to perform attention fusion on the information within and between meta-paths to generate graph node embeddings.

[0021] Step 4, use the contrastive learning method to establish a contrastive loss function based on the node embeddings from the two perspectives, and train to generate a cross-language source code representation model.

[0022] Step 4.1, establish positive and negative sample sets according to the correlation between nodes.

[0023] Step 4.2, establish a contrastive loss function and a model optimization objective function based on the node embeddings from the two perspectives, and use the backpropagation algorithm for optimization to train and generate a cross-language source code representation model.

[0024] Step 4.3, after contrastive learning, select the node embeddings from the serialized perspective as the final source code representation vectors.

[0025] Beneficial effects

[0026] Compared with the tree-based code representation method, the present invention incorporates language-specific information into the Transformer structure by generating dynamic parameters according to the language type of the input code snippet and further generating the weight matrix of the learnable parameters in the self-attention module, ensuring the learning of language-independent information and language-specific information of the source code and improving the cross-language source code representation effect of the model.

[0027] Compared with the graph-based code representation method, the present invention constructs a model by combining the local aggregation ability of GCN and the global modeling ability of meta-paths, using multi-view graph node embeddings, ensuring the learning of local and global information of the heterogeneous graph of source code features, avoiding the influence of the over-smoothing problem of the model in the graph representation learning process, and improving the cross-language source code representation effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the multi-view contrastive learning cross-language source code representation method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] To better illustrate the purpose and advantages of the present invention, the following further elaborates on the implementation manner of the method of the present invention in combination with examples.

[0030] The experimental data comes from Python, Ruby, Javascript, and Go language codes in the CodeSearchNet public dataset.

[0031] Table 1. CodeSearchNet Dataset

[0032]

[0033] The specific process is as follows:

[0034] Step 1, initialize the learnable parameters of the Transformer self-attention module according to the programming language type using the meta-learning method.

[0035] Specifically, for a given source code snippet X i and its corresponding language type t i , first pass through the language embedding layer to obtain the vector T i ∈R dT . Then, scale the dimension of T i through the projection layer.

[0036]

[0037] P i = Projection(T i ) (2)

[0038] Among them, Projection(·) and P i ∈R dp respectively represent the projection layer and the projection language type embedding.

[0039] After obtaining the projection language type embedding P i , the weight matrix W of the learnable parameters in the attention module is output through the parameter generator λ ∈R d×d .

[0040]

[0041] The method decomposes the representation of the weights in a way similar to singular value decomposition, and projects the original vector into the target matrix dimension space with fewer parameters in the generator. Apply the diagonalization operation to transform the vector P i into a diagonal matrix, and then define two projection matrices and to project the diagonalized embedding matrix into the space of the target weight matrix.

[0042]

[0043] Among them, diag(·) is a non-parametric diagonal operation, M λ and M′ λ are the learnable parameters in the generator .

[0044] Step 2: Construct a code feature heterogeneous graph containing code syntax, control, and dependency relationships for the source code fragment, and use GCN to aggregate the heterogeneous graph to obtain graph node embeddings.

[0045] Step 2.1: Use the open-source tools Joern, Soot, and ANTLR to generate AST, PDG, and CFG representations of code in different languages.

[0046] Step 2.2: Expand based on the AST and incorporate the PDG and CFG. Traverse the edge sets of the PDG and CFG graphs, and add edges between the corresponding nodes in the AST according to the index information of the two end nodes to generate a heterogeneous graph with two types of structural information, control and dependency.

[0047] Step 2.3: Perform node feature transformation to unify the feature dimensions of different types of nodes in the heterogeneous graph.

[0048] Unify the node feature dimensions by creating a mapping matrix h i for each node type in the heterogeneous graph.

[0049]

[0050] where is the original feature of node i, is the type of node i, is the original feature dimension of the type node, is the mapping matrix of the type node, is the corresponding bias vector, d is the dimension of the target mapping space, and σ is the activation function.

[0051] Step 2.4, Use GCN to aggregate the heterogeneous graph to obtain the node embedding from the graph perspective.

[0052]

[0053] where L g is the number of layers of the graph neural network, represents the nearest neighbor of node i under relationship r, is the feature representation of the nearest neighbor node j of node i under relationship r, represents the weight matrix related to relationship r, represents the weight matrix related to the feature of node i itself, c i,r is a normalization constant, h i is the projection feature matrix obtained in Step 2.3, represents the feature vector obtained by updating node i after passing through l layers of graph convolutional neural networks. After passing through L g layers of network processing, the node embedding of node i from the graph perspective is obtained

[0054] Step 3, Aggregate the node co-hop neighborhood information according to the meta-path to obtain the sequential representation of the graph nodes, and use Transformer and hierarchical attention mechanism to generate the node embedding from the sequence perspective.

[0055] Step 3.1, Use the meta-path perception strategy to generate the representation vectors of graph nodes under each meta-path. First, obtain the different-hop neighbor nodes of nodes under different meta-paths.

[0056] For node i, define as the k-hop neighbor of node i under the meta-path where d(i, j) represents the shortest path between node i and node j, and the 0-hop neighbor is the node itself, N 0 (i) = {i}.

[0057] Then, regard the neighbors from the same hop as a group, and perform information aggregation within the group, and splice to obtain the representation vector of the node under the meta-path

[0058]

[0059] Among them, is the adjacency matrix under the meta-path The following is the adjacency matrix, is the meta-path The following is the sequence representation of the k-hop neighbors under the meta-path. H is the projection feature matrix obtained in step 2.3, D is The degree matrix of, n is the node number, and d is the sequence dimension. Assuming that the maximum number of hops is K, for node i, the representations of the neighbor nodes at each hop under the meta-path can be obtained where k ∈ 0, 1, 2, …, K, and after splicing, the serialized representation vector of node i under the meta-path is obtained

[0060] Step 3.2: Use Transformer to encode the node representation vectors to further explore the semantic interactions between different-hop neighborhoods of nodes under the same meta-path.

[0061]

[0062] Among them is the learnable mapping matrix, is the projection representation with dimension d m . Input into the multi-layer encoder module to explore the semantic relationships between different-hop neighbors. Each encoder module consists of a multi-head attention layer (MSA), a feed-forward layer (FFN), and a residual normalization layer (LN).

[0063]

[0064] Among them is the intermediate output of the input sequence after passing through the residual normalization layer and the multi-head attention layer, is the output of the multi-head attention layer The node representation obtained after passing through the residual normalization layer and the feed-forward neural network layer. l = 1, 2, …, L, where L is the number of encoders. After L-layer encoding, a meta-path-based node representation containing richer semantic information can be obtained

[0065] Step 3.3: Use the hierarchical attention mechanism to perform attention fusion on intra-meta-path and inter-meta-path information to generate graph node embeddings.

[0066] Considering the intra-meta-path information, by calculating the correlation between neighbor nodes and the node itself, the weights of neighbor node information at different hop numbers for the embedding representation of a meta-path are obtained

[0067]

[0068] Among them, is a learnable parameter matrix, represents the k-hop representation of node i under the meta-path Furthermore, the information aggregation representation of neighbor nodes with different hop numbers is realized

[0069]

[0070] Considering the information between meta-paths, since different meta-paths express different semantics, their contributions to the final representation of nodes are also different in different tasks or different datasets. It is necessary to calculate the weights corresponding to the representation vectors of the same node under different meta-paths

[0071]

[0072] where Φ is the set of meta-paths, and are the learnable parameter matrices of the meta-path , tanh is the activation function, and finally the graph node representation from the serialization perspective can be obtained

[0073]

[0074] Step 4: Use the contrastive learning method to establish a contrastive loss function based on the node embeddings from two perspectives, and train to generate a cross-language source code representation model.

[0075] Step 4.1: Establish positive and negative sample sets according to the correlation between nodes.

[0076] In a heterogeneous graph, different meta-paths represent different semantic correlations. If there are multiple meta-path instances between two nodes, it can be proved that there is a high correlation between the nodes. The positive and negative samples are divided by counting the number of meta-paths between nodes.

[0077]

[0078] Among them, C i (j) represents the number of meta-paths between node i and node j. Set a threshold θ. If C i (j) ≥ θ, add the node pair (i, j) to the positive sample set P i of node i, otherwise add it to the negative sample set N i .

[0079] Step 4.2, establish a contrastive loss function and a model optimization objective function based on the node embeddings from two perspectives, and use the backpropagation algorithm for optimization to train and generate a cross - language source code representation model.

[0080]

[0081] Among them, sim(i, j) represents the cosine similarity between vector i and vector j, and τ is the temperature coefficient. represents the contrastive loss from the graph - perspective node embedding to the sequential - perspective node embedding. represents the contrastive loss from the sequential perspective to the graph perspective, L is the overall objective function, λ is the balance coefficient between the two perspectives, and V represents the set of all graph nodes.

[0082] Step 4.3, after contrastive learning, select the node embeddings from the sequential perspective as the final source code representation vectors.

Claims

1. Multi-perspective comparative learning cross-language source code representation method, characterized by The method comprises the following steps: Step 1: Use meta-learning to initialize the learnable parameters of the Transformer self-attention module according to the programming language type, and introduce language-specific information into the model; Step 2: construct a code feature heterogeneous graph containing code syntax, control, and dependency relationships for the source code snippet, and use GCN to aggregate the heterogeneous graph to obtain node embedding from the graph perspective; Step 3: Aggregate the node same-hop neighborhood information according to the meta-path to obtain the serialized representation of the graph node, and use Transformer and hierarchical attention mechanism to generate graph node embedding: First, use the meta-path perception strategy to generate the representation vector of the graph node under each meta-path; then use Transformer to encode the node representation vector, and further explore the semantic interaction between nodes in different hop neighborhoods under the same meta-path; finally, use the hierarchical attention mechanism to fuse the information within and between meta-paths to generate node embedding from a serialized perspective; Step 4: Apply the contrastive learning method, establish a contrastive loss function based on the node embeddings from two perspectives, and train to generate a cross-language source code representation model.

2. The multi-perspective contrast learning cross-language source code representation method according to claim 1, characterized in that: In step 1, the meta-learning method is used to generate dynamic parameters according to the programming language type of the input code snippet, and further initialize the learnable parameters in the Transformer self-attention module, thereby introducing language-specific information learning.

3. The multi-perspective contrast learning cross-language source code representation method according to claim 1, characterized in that: In step 2, a code feature heterogeneous graph including code syntax, control and dependency relationships is constructed for the source code fragment. The code feature heterogeneous graph is constructed by supplementing the control information and dependency relationships contained in the control flow graph (CPG) and the program dependency graph (PDG) on the basis of the abstract syntax tree (AST).

4. The multi-perspective contrast learning cross-language source code representation method according to claim 1, characterized in that: In step 3, the meta-path awareness strategy is used to generate the representation vector of the graph node under each meta-path. First, the neighbor nodes with different hops of the node under different meta-paths are obtained, and then the neighbors from the same hop are regarded as a group and information aggregation is performed within the group. in, is the meta path The adjacency matrix under The degree matrix of is the meta path The sequence representation of the k-hop neighbors under the condition, H is the projection feature matrix obtained by node feature transformation in step 2. Assuming that the maximum number of hops is K, for node i, we can obtain The representation of each hop neighbor node under Splice to get node i in the meta path The serialized representation vector is 5. The multi-perspective contrast learning cross-language source code representation method according to claim 1, characterized in that: In step 3, a hierarchical attention mechanism is used to fuse the information within and between meta-paths. Considering the information within the meta-path, the importance of neighbor nodes with different hop numbers to the embedded representation of a meta-path is explored by calculating the correlation between neighbor nodes and the node itself, and information aggregation between neighbor nodes with different hop numbers is achieved. Considering the information between meta-paths, since different meta-paths express different semantics, the representation vectors of the same node under different meta-paths are aggregated to obtain the graph node representation from a serialization perspective.

6. The cross-language source code representation method for multi-view contrastive learning according to claim 1, characterized in that: In step 4, the contrastive learning method is applied to establish a contrastive loss function based on the node embeddings from two perspectives, and train and generate a cross-language source code representation model.

Citation Information

Patent Citations

  • Heterogeneous graph embedding learning method based on attention mechanism

    CN113095439A

  • Code vulnerability detection method based on graph contrast learning

    CN116204877A

  • Self-supervised heterogeneous graph representation learning method capable of resisting excessive smoothness

    CN117473124A

  • Cross-language code similarity detection system and method based on graph attention network

    CN118227139A

  • Vulnerability detection method based on source code multi-scale feature extraction and comparative learning

    CN119357974A