Code pre-training method based on Adapter network and comparative learning

By introducing Adapter network and comparative learning in the code pre-training method, the problem of programming language knowledge forgetting and insufficient resource support capabilities in the prior art is solved, and better multilingual support and model performance in low resource environments is achieved.

CN120029600AActive Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411961870.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-23
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing code pre-training methods have catastrophic forgetting problems in code mix corpus of different programming languages ​​and have weak support for low-resource programming languages, resulting in poor performance of models in adding new programming language support and low-resource environments.

Method used

Using the code pre-training method based on Adapter network and contrast learning, we use multiple Transformer layers and programming language-specific Adapter networks to reduce knowledge forgetting, and build positive and negative examples for low-resource programming languages, and use a joint training model for classification cost and contrast learning cost.

Benefits of technology

It effectively reduces the catastrophic forgetting problems of knowledge of different programming languages, supports rapid training of new programming languages, and enhances the capabilities of low-resource programming language models, improving the overall performance of the model in multi-language environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029600A_ABST
    Figure CN120029600A_ABST
Patent Text Reader

Abstract

The invention discloses a code pre-training method based on an Adapter network and comparative learning. The method comprises the following steps: constructing a code pre-training model based on the Adapter network, obtaining a training instance xd in a code corpus, training the code pre-training model based on the Adapter network, and obtaining the Adapter network with different programming language knowledge; constructing a low-resource programming language model; obtaining an instance xc in the code corpus, respectively constructing a corresponding positive example and a corresponding negative example, and jointly training a low-resource programming language model by adopting classification cost and contrast learning cost; and training to obtain a final code pre-training model, and outputting a code pre-training result based on the final code pre-training model. According to the method, the problem of disastrous forgetting of existing knowledge can be reduced while code corpora of multiple programming languages are pre-trained, training of a new programming language is supported, and the model capability under a low-resource programming language can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code pre-training, and in particular to a code pre-training method based on an adapter network and contrastive learning. Background Art

[0002] The purpose of the code generation task is to automatically generate executable code based on a high-level description of the input (usually a natural language that describes human intent). The automatically generated code can be a complete application, system component, or just some specific modules, functions, or code snippets. Automatic code generation can reduce repetitive programming tasks and help developers focus on more innovative and complex business logic. By automatically generating common code templates or snippets, the time for manual writing is reduced, thereby greatly improving the development efficiency of programmers.

[0003] Automatic code generation methods can be roughly divided into the following two categories: 1) Methods based on sequence-to-sequence networks: Based on the encoder-decoder network architecture and the integrated attention mechanism, the correspondence between natural language intent and code is learned from the training data set. Specifically, the encoder is responsible for encoding the natural language intent input into a fixed-length vector representation, which represents the abstract representation information of the natural language intent. The decoder decodes and generates executable code step by step based on the encoder's vector representation. 2) Methods based on large pre-trained models: This type of method is divided into two stages. The pre-training stage trains a code pre-training model with general code knowledge based on a large number of code corpora in different programming languages. The fine-tuning stage uses the model obtained in the previous stage for fine-tuning of the automatic code generation task.

[0004] Currently, methods based on sequence-to-sequence networks are only trained on a small amount of code generation task data. Although they can learn some mapping relationships, they lack a general understanding of the code itself, resulting in performance that is generally lower than that of methods based on large pre-trained models. In addition, the lack of manually labeled training data on code generation often leads to overfitting problems in models based on sequence-to-sequence networks.

[0005] However, the large-scale pre-trained model-based approach has the following two shortcomings: First, the large-scale pre-trained model-based approach mixes the corpora of different code language fields into the model for training, and then uses the pre-trained model to fine-tune the automatic code generation task. Therefore, after training the code pre-trained model, if the model wants to add support for a new code language, the entire model needs to be retrained using the new code language; moreover, integrating different languages ​​at the same time will lead to catastrophic forgetting problems, and the number of pre-trained model parameters is generally very large, and the cost of retraining is high. Second, for some low-resource programming languages, the available code corpus is very small, so compared to the high-resource pre-trained models trained with abundant data, the model capabilities of low-resource programming languages ​​are weaker.

[0006] When existing methods pre-train Transformer models in a mixed corpus of codes in different programming languages, since Transformer models generally contain a large number of parameters, the amount of code corpora in different programming languages ​​varies greatly. For example, the code of the Ruby programming language accounts for a very small proportion of the entire corpus, resulting in existing methods often forgetting previously learned programming language knowledge during pre-training, resulting in the problem of catastrophic forgetting of existing knowledge. In addition, when it is necessary to support a new programming language, the entire model needs to be retrained, which consumes a lot of resources and time for training. Summary of the invention

[0007] In order to overcome the defects and shortcomings of the prior art, the present invention provides a code pre-training method based on Adapter network and contrastive learning. The present invention constructs a code pre-training model based on Adapter network, and obtains Adapter network with knowledge of different programming languages ​​based on a large number of different programming language corpora. The Adapter network can reduce the problem of catastrophic forgetting of existing knowledge while pre-training code corpora of multiple programming languages, and supports the training of new programming languages; constructs positive and negative samples for data of low-resource programming languages, and adopts classification cost and contrastive learning cost to jointly train the low-resource programming language model, which can enhance the model capability under low-resource programming languages.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention provides a code pre-training method based on an adapter network and contrastive learning, comprising the following steps:

[0010] Build a code pre-training model based on the Adapter network, including: multiple Transformer layers, Adapter networks corresponding to different programming languages, a fusion layer, and a classification layer. The multiple Transformer layers learn the general knowledge of programming languages, the Adapter networks learn the domain knowledge of the current programming language, the fusion layer fuses the general knowledge and the domain knowledge of different programming languages, and the classification layer calculates the prediction results of masked languages;

[0011] Obtain the training instance x in the code corpus d , and train the code pre-training model based on the Adapter network to obtain an Adapter network with knowledge of different programming languages;

[0012] Build a low-resource programming language model, including an input layer, an encoding layer, a pooling layer, and a classification layer. The input layer inputs the instances in the code corpus, the encoding layer calculates the semantic representation of the instances in the code corpus, the encoding layer is initialized by the pre-trained Adapter network corresponding to the programming language knowledge, the pooling layer calculates the semantic vector representation of the instances, and the classification layer calculates the prediction results for the masked words;

[0013] Obtain the instance x in the code corpus c , respectively construct corresponding positive and negative examples, and jointly train the low-resource programming language model using the classification cost and the contrastive learning cost;

[0014] Train to obtain the final code pre-training model, and output the code pre-training results based on the final code pre-training model.

[0015] As a preferred technical solution, the general knowledge of the programming language includes: data structures and algorithms, as well as programming paradigms and control flow structures.

[0016] As a preferred technical solution, obtain the training instance x in the code corpus d , and train the code pre-training model based on the Adapter network, specifically including:

[0017] The Transformer layer is initialized based on the parameters of the code pre-training model CodeBERT;

[0018] Input the training instance x d , and obtain the corresponding output after passing through multiple Transformer layers, specifically expressed as:

[0019]

[0020] Among them, and respectively represent the k-th and 12th Transformer layers of the code pre-training model CodeBERT, and Represents the outputs corresponding to the kth and 12th Transformer layers of the code pre-training model CodeBERT;

[0021] Each layer of the Adapter network consists of a multi-layer feedforward neural network and a Transformer layer;

[0022] Given a programming language p and an instance x d The initial semantic matrix representation of And the output of the first Transformer layer of the code pre-trained model CodeBERT The Adapter network is transformed as follows:

[0023]

[0024] Among them, FFN 1 FFN 2 and Represents the 1st, 2nd and Nth p Multi-layer feed-forward neural network in the Adapter layer, The first, second and Nth p The Transformer layer in the Adapter layer, and The first, second and Nth p The output of the Adapter layer, [ ] represents the matrix concatenation operation;

[0025] The fusion layer includes a concatenation layer and an MLP layer. The fusion layer calculates its final semantic vector representation h as follows: d,p :

[0026]

[0027] Where [ ] represents the matrix concatenation operation of the concatenation layer, and MLP represents the MLP layer;

[0028] The classification layer consists of a fully connected layer and a softmax transformation to obtain the instance x in the code corpus of a given programming language p d The final semantic vector representation h d,p After that, the classification layer calculates the prediction result of the mask language Specifically expressed as:

[0029]

[0030] Among them, W 1 and b 1 are the parameters in the fully connected layer, It is a v-dimensional vector, where v represents the size of the vocabulary during pre-training.

[0031] As a preferred technical solution, when training the code pre-training model based on the Adapter network, given the code corpus D c The example in (x d ,y d ), the cross entropy cost function of the masked language model is defined as follows:

[0032]

[0033] Among them, θ c represents the parameter set of the entire model, Represents the expected value of the predicted result with respect to the true label.

[0034] As a preferred technical solution, the encoding layer calculates the semantic representation of the instance in the code corpus, which is specifically expressed as:

[0035] H c =Adapter p (x c )

[0036]

[0037] Among them, H c , and For instance x c , positive example and negative examples The semantic matrix representation of , each column in the matrix corresponds to a word in the instance, Adapter p represents the coding layer;

[0038] The pooling layer calculates the semantic vector representation of the instance, specifically:

[0039] h c =AvgPooling(H c )

[0040]

[0041] Among them, h c , and For instance x c , positive example and negative examples The semantic vector representation of , AvgPooling is the average pooling operation;

[0042] The classification layer includes a fully connected layer and a softmax transformation.c The semantic vector representation h c After that, the classification layer of the masked language model calculates the prediction result for the masked word Specifically expressed as:

[0043]

[0044] Among them, W 2 and b 2 are the parameters of the fully connected layer, is a v-dimensional vector, where v is the size of the vocabulary in the corpus.

[0045] As a preferred technical solution, the classification cost adopts the cross entropy cost function, which is specifically expressed as:

[0046]

[0047] Among them, θ′ c represents the parameter set of the connective classification model, y c For instance x c The true category to which the masked word belongs, It represents the expected value of the predicted word category with respect to the category of the real word.

[0048] As a preferred technical solution, the contrastive learning cost is expressed as:

[0049]

[0050] Among them, h c , and For instance x c , positive example and negative examples The semantic vector representation of , sim() is used to measure the cosine distance between two vectors, || || represents the 2-norm of the vector, and T represents the transpose of the vector.

[0051] As a preferred technical solution, the total cost of the model is obtained by linearly summing the classification cost and the contrastive learning cost, which is expressed as:

[0052] L(θ′ c )=L MLM (θ′ c )+λL cl (θ′ c )

[0053] Among them, λ is the weight coefficient, which is used to adjust the importance of pre-training cost and comparative learning cost, L MLM (θ′ c ) represents the classification cost, L cl (θ′c ) represents the contrastive learning cost.

[0054] The present invention also provides a code pre-training system based on Adapter network and contrastive learning, including: a code pre-training model construction module based on Adapter network, a code pre-training model training module based on Adapter network, a low-resource programming language model construction module, a low-resource programming language model training module, and a code pre-training result output module;

[0055] The Adapter network-based code pre-training model construction module is used to construct a code pre-training model based on the Adapter network, including: multiple Transformer layers, Adapter networks corresponding to different programming languages, a fusion layer and a classification layer, the multiple Transformer layers learn general knowledge of programming languages, the Adapter network learns the domain knowledge of the current programming language, the fusion layer fuses general knowledge and domain knowledge of different programming languages, and the classification layer calculates the prediction result of the mask language;

[0056] The code pre-training model training module based on the Adapter network is used to obtain the training instance x in the code corpus. d , train the code pre-training model based on the Adapter network to obtain the Adapter network with knowledge of different programming languages;

[0057] The low-resource programming language model construction module is used to construct a low-resource programming language model, including an input layer, an encoding layer, a pooling layer and a classification layer, the input layer inputs instances in the code corpus, the encoding layer calculates semantic representations of instances in the code corpus, the encoding layer is initialized by a pre-trained Adapter network corresponding to programming language knowledge, the pooling layer calculates semantic vector representations of instances, and the classification layer calculates prediction results for masked words;

[0058] The low-resource programming language model training module is used to obtain instances x in the code corpus. c , respectively construct corresponding positive and negative examples, and use classification cost and contrastive learning cost to jointly train the low-resource programming language model;

[0059] The code pre-training result output module is used to output the code pre-training result based on the final code pre-training model.

[0060] The present invention also provides a computer-readable storage medium storing a program, wherein when the program is executed by a processor, the code pre-training method based on an adapter network and contrastive learning is implemented as described above.

[0061] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0062] The present invention constructs a code pre-training model based on an Adapter network, and obtains an Adapter network with knowledge of different programming languages ​​based on a large number of different programming language corpora. The Adapter network can reduce the problem of catastrophic forgetting of existing knowledge while pre-training code corpora of multiple programming languages, and support the training of new programming languages. Positive and negative samples are constructed for data of low-resource programming languages, and classification cost and contrastive learning cost are used to jointly train the low-resource programming language model, which can enhance the model capability under low-resource programming languages. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic diagram of the overall implementation architecture of the code pre-training method based on Adapter network and contrastive learning of the present invention;

[0064] Figure 2 It is a flowchart of the code pre-training method based on Adapter network and contrastive learning of the present invention;

[0065] Figure 3 This is a schematic diagram of the overall network architecture of the code pre-training model based on the Adapter network of the present invention;

[0066] Figure 4 It is a schematic diagram of the implementation process of the low-resource programming language model enhancement stage of the present invention. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0068] Example 1

[0069] like Figure 1 As shown, this embodiment provides a code pre-training method based on Adapter network and contrastive learning, which is divided into two stages:

[0070] 1) Multi-programming language code pre-training stage: In this stage, a code pre-training model based on an Adapter network is built based on a large number of different programming language corpora. The model designs a general Transformer layer as the backbone network for each code language, and uses the existing code pre-training model for initialization to learn general knowledge. An Adapter network is built for each programming language to learn the domain knowledge of different programming languages, thereby maintaining the independence of the knowledge of each programming language, reducing the problem of catastrophic forgetting, and facilitating the pre-training of subsequent new programming languages. After training the model based on the Adapter network, if a new programming language needs to be supported, such as C language, only the Adapter network of the C language needs to be added, and its Adapter network is trained using the C language code corpus data. Other code domains will not be affected.

[0071] 2) Low-resource programming language enhancement stage: This stage aims to enhance the adapter corresponding to the low-resource programming language pre-trained in the previous stage. Specifically, it constructs positive and negative samples for the data of the low-resource programming language, uses contrastive learning to shorten the distance between positive samples and increase the distance between negative samples, and learns better code representations of low-resource programming languages, thereby enhancing the capabilities of the adapter network of low-resource programming languages. The present invention can not only avoid the catastrophic forgetting problem of existing knowledge, but also support the training of new programming languages ​​at a very low training cost, and enhance the capabilities of the model on low-resource programming languages.

[0072] like Figure 2 As shown, the method specifically comprises the following steps:

[0073] S1: Build a code pre-training model based on the Adapter network. The model architecture includes: multiple Transformer layers, Adapter networks corresponding to different programming languages, fusion layers, and classification layers.

[0074] Specifically, multiple Transformer layers are used to learn common knowledge of all programming languages. In this embodiment, 12 Transformer layers are preferred because data of all programming languages ​​will pass through the Transformer layer. Therefore, the Transformer layer updates parameters based on the data of all programming languages. The parameter update allows the model to learn knowledge, so the knowledge of all programming languages ​​is integrated into this Transformer layer.

[0075] The Adapter layer is specific to different programming languages ​​and is used to learn the domain knowledge of the current programming language. Since the Adapter layer only updates the corresponding programming language data, the domain knowledge of the corresponding programming language is trained.

[0076] The fusion layer is used to integrate general knowledge and domain knowledge of different programming languages;

[0077] The classification layer is used to calculate the prediction results of the masked language;

[0078] In this embodiment, the general knowledge of programming languages ​​refers to some basic concepts and knowledge that all programming languages ​​have, such as data structures and algorithms, programming paradigms, and control flow structures in each language, as follows:

[0079] Data structures and algorithms: Data structures: arrays, linked lists, stacks, queues, trees, graphs, hash tables, etc. Algorithms: sorting, searching, recursion, dynamic programming, graph algorithms, greedy algorithms, etc.

[0080] Programming paradigm: For example, Java, C++, and Python are all object-oriented programming languages.

[0081] Control flow structure: For example, all languages ​​have control statements similar to if-else;

[0082] S2: Figure 3 As shown, get the training instance x d , train the code pre-training model based on the Adapter network to obtain the Adapter network with knowledge of different programming languages, including:

[0083] The Transformer layer is initialized based on the parameters of the code pre-trained model CodeBERT, and the parameters remain unchanged in the subsequent training;

[0084] Input training instance x d , after multiple Transformer layers, the corresponding output is obtained, which is specifically expressed as:

[0085]

[0086] in, and They represent the kth and 12th Transformer layers of the code pre-training model CodeBERT, respectively. and Represents the outputs corresponding to the kth and 12th Transformer layers of the code pre-training model CodeBERT, Can be used as training instance x d The initial semantic matrix representation of k is a hyperparameter of the model and can be set manually;

[0087] In this embodiment, the training instance x dIt can be a piece of code, such as the Python code def [mask word] (): print ("Hello"), the model needs to predict what the mask word should be filled in, for example, it should predict hello;

[0088] For different programming languages, the Adapter network can set different numbers of Adapter layers. The more code corpus a programming language has, the more knowledge it contains, and the more parameters of Adapter layers are needed to learn. This adaptive Adapter layer design can also allocate the best number of Adapter layers for each programming language, thereby achieving better results with as few parameters as possible.

[0089] For programming language p, set N p Adapter layer, N p It is a hyperparameter of the Adapter network. In practical applications, it can be manually specified according to the number of code corpora of programming languages. For example, if pyton has 10,000 data, java has 5,000 data, and c++ has 5,000 data, then the proportion of the total data volume they occupy is 0.5, 0.25, and 0.25. If a maximum number of Adapter layers is manually defined, for example, 10 layers, then the number of layers of these three languages ​​can be set to 0.5*10=5 layers and 0.25*10=2.5 ​​(rounded up, can be set to 3 layers).

[0090] In this embodiment, each layer of the Adapter network consists of a multi-layer feed-forward neural network (FFN) and a Transformer layer. The Adapter network is based on the output of the first Transformer layer of the code pre-training model CodeBERT. and the initial semantic matrix representation of the Adapter network These semantic matrix representations (features) are further transformed to obtain semantic matrix representations (features) that are more suitable for each programming language. Compared with directly adjusting the parameters in the pre-trained encoding layer CodeBERT, the Adapter network adjusts the pre-trained model from another perspective, making it well suited for pre-training tasks in different programming languages.

[0091] Specifically, given a programming language p and an instance x d The initial semantic matrix representation of And the output of the first Transformer layer of the code pre-trained model CodeBERT The Adapter network is transformed as follows:

[0092]

[0093] Among them, FFN1 FFN 2 and Represents the 1st, 2nd and Nth p Multi-layer feed-forward neural network in the Adapter layer, The first, second and Nth p The Transformer layer in the Adapter layer, and The first, second and Nth p The output of the Adapter layer, [] represents the matrix concatenation operation, is the initial input of the Adapter network, which can be randomly initialized or initialized to a zero matrix. The multi-layer feedforward neural network transforms the input vector to a smaller dimension, so that the parameters in the Adapter network are relatively small. The final output of the Adapter network is Can be seen as instance x d The transformed semantic matrix representation.

[0094] In this embodiment, the programming language p includes but is not limited to Java, C++, Python, and ruby;

[0095] In this embodiment, the knowledge of different programming languages ​​is focused on learning the knowledge of their own programming languages ​​through different Adapter layers, and the common knowledge between different programming languages ​​is learned through a shared Transformer layer, thereby avoiding the problem of catastrophic forgetting and being able to quickly support pre-training of new programming languages. Removing the Adapter network of a certain programming language will not affect other programming languages.

[0096] When getting instance x d Output of the 12th Transformer layer in the code pre-trained model CodeBERT And the transformed semantic matrix representation After that, these two representations are input into the fusion layer for fusion. The fusion layer includes a concatenation layer and an MLP layer. Specifically, the fusion layer calculates its final semantic vector representation h as follows: d,p :

[0097]

[0098] Where [ ] represents the matrix concatenation operation of the concatenation layer, and MLP represents the MLP layer;

[0099] The classification layer consists of a fully connected layer and a softmax transformation to obtain the instance x in the code corpus of a given programming language p d The final semantic vector representation hd,p After that, the classification layer calculates the prediction result of the mask language As shown below:

[0100]

[0101] Among them, W 1 and b 1 are the parameters in the fully connected layer, It is a v-dimensional vector, where v represents the size of the vocabulary during pre-training.

[0102] During training, given the code corpus D c The example in (x d ,y d ), the cross entropy cost function of the masked language model is defined as follows:

[0103]

[0104] Among them, θ c represents the parameter set of the entire model, Represents instance x d The categories of all masked words in , Indicates the expected value of the predicted result with respect to the true label. Mask language refers to the way the model is trained in a masked way, such as instance x d This is a piece of Python code "def[mask word]():print("hello"), the classification layer predicts the corresponding mask word, and ideally it will output "hello". After entering the Python adapter network, the output of the classification layer is the probability of various words in a vocabulary, and the word with the largest probability is taken as the word corresponding to the mask word. If the classification layer predicts incorrectly, a loss will occur. Driven by the loss, continuous training will make the model's prediction more and more accurate, closer to the prediction of "hello".

[0105] The code pre-training model based on the Adapter network is pre-trained on the code corpus of a large number of programming languages ​​to obtain the Adapter network with knowledge of different programming languages;

[0106] S3: Build a low-resource programming language model. The model architecture includes input layer, encoding layer, pooling layer and classification layer.

[0107] S4: Figure 4 As shown, the classification cost and contrastive learning cost are used to jointly train the low-resource programming language model;

[0108] For some low-resource programming languages, the available code corpus is very small. Therefore, compared with the high-resource Adapter network trained with abundant data, the Adapter network of the low-resource programming language is weaker and it is difficult to learn enough knowledge of the low-resource programming language. This embodiment constructs positive examples (similar examples) and negative examples (dissimilar examples) of a given instance, and through learning, makes the distance between it and the positive example in the semantic space closer, and the distance between it and the negative example in the semantic space farther, so that the model has a stronger ability to recognize domain knowledge.

[0109] First, we mine positive and negative samples of low-resource programming languages ​​by data augmentation. In the input layer, for an instance x in a given code corpus, c , respectively construct the corresponding positive examples and negative examples

[0110] For positive examples The construction of the code is different from that of natural language. Slight changes to the code may greatly change the original meaning of the code. Therefore, data enhancement is performed on the code based on the semantic invariance rule. Specifically, the abstract syntax tree of the code is first extracted, and a variable in the abstract syntax tree (a variable customized by the user when writing the program) is randomly named, and the corresponding variables in other locations in the entire code are also named in the same way, so as to ensure that the function and semantics of the entire code are consistent. Then, the loop statement is rewritten, such as rewriting the for loop into a while statement with the same semantics, which can also ensure that the code semantics are consistent. Based on the above two semantically consistent data enhancements, we can perform data enhancement for instance x. c Generate positive examples in various forms.

[0111] The construction of negative examples can be divided into two cases: c Different other instances can be used as general negative examples. You can also design some instances that are more difficult to distinguish, such as instances with similar code functions, as high-quality negative examples.

[0112] Take instance x c 、The positive example and negative examples Combined into a training instance, denoted as

[0113] Given a low-resource code instance x c 、The positive example and negative examples Coding layer Adapter p For computing the semantic representation of these instances, for low-resource programming languages ​​​​p, the encoding layer Adapter pInitialized by the pre-trained Adapter network corresponding to the programming language knowledge, specifically, the input instance x c 、The positive example and negative examples Through the encoding layer Adapter p After calculation, we get:

[0114] H c =Adapter p (x c )

[0115]

[0116] Among them, H c , and For instance x c , positive example and negative examples The semantic matrix representation of is , where each column in the matrix corresponds to a word in the instance.

[0117] The pooling layer only includes the average pooling operation to calculate the semantic vector representation of the instance, as shown below:

[0118] h c =AvgPooling(H c )

[0119]

[0120] Among them, h c , and For instance x c , positive example and negative examples The semantic vector representation of , AvgPooling is the average pooling operation;

[0121] The classification layer consists of a fully connected layer and a softmax transformation. c The semantic vector representation h c After that, the prediction result of the masked word can be obtained through the classification layer calculation of the masked language model (MLM). As shown below:

[0122]

[0123] Among them, W 2 and b 2 are the parameters of the fully connected layer, is a v-dimensional vector, where v is the size of the vocabulary in the corpus.

[0124] Specifically, given a code corpus D c The training examples in The cross entropy cost function is defined as follows:

[0125]

[0126] Among them, θ′ c represents the parameter set of the connective classification model, y c For instance x c The true category to which the masked word belongs, E yc [] represents the expected value of the predicted word category with respect to the category of the real word;

[0127] In addition to the cross entropy classification cost, the cost based on contrastive learning is also used, and the contrastive learning cost function is further defined as follows:

[0128]

[0129] Among them, h c , and For instance x c , positive example and negative examples The semantic vector representation of , sim() is used to measure the cosine distance between two vectors, || || represents the 2-norm of the vector, T represents the transpose of the vector, and minimizing the contrastive learning cost can make the instances in the training data set closer to their positive examples in the semantic space, while the distance between them and the negative examples is relatively far, which is conducive to distinguishing codes with similar semantics. Finally, the total cost of the model is the linear sum of the classification cost and the contrastive learning cost, as follows:

[0130] L(θ′ c )=L MLM (θ′ c )+λL cl (θ′ c )

[0131] Among them, λ is the weight coefficient, which is used to adjust the importance of pre-training cost and comparative learning cost.

[0132] The low-resource programming language model enhancement method based on contrastive learning can further enhance the effect of the low-resource programming language Adapter network. The classification layer predicts the probabilities of different words. The classification cost is to calculate the difference between the predicted words and the real words. The goal of contrastive learning cost training is to make the code examples more similar to the positive examples and less similar to the negative examples. In this way, the low-resource Adapter is equivalent to training more data (because one code can generate multiple positive examples, there will be more training data), and it can also distinguish different low-resource codes, which is equivalent to enhancing the Adapter's understanding of such low resources.

[0133] S5: Train to obtain a final code pre-training model, and output a code pre-training result based on the final code pre-training model.

[0134] Example 2

[0135] This embodiment provides a code pre-training system based on an adapter network and contrastive learning, which is used to implement the code pre-training method based on an adapter network and contrastive learning in Embodiment 1, including: a code pre-training model construction module based on an adapter network, a code pre-training model training module based on an adapter network, a low-resource programming language model construction module, a low-resource programming language model training module, and a code pre-training result output module;

[0136] In this embodiment, the code pre-training model construction module based on the Adapter network is used to construct a code pre-training model based on the Adapter network, including: multiple Transformer layers, Adapter networks corresponding to different programming languages, a fusion layer and a classification layer, the multiple Transformer layers learn the general knowledge of the programming language, the Adapter network learns the domain knowledge of the current programming language, the fusion layer fuses the general knowledge and the domain knowledge of different programming languages, and the classification layer calculates the prediction result of the mask language;

[0137] In this embodiment, the code pre-training model training module based on the Adapter network is used to obtain the training instance x in the code corpus. d , train the code pre-training model based on the Adapter network to obtain the Adapter network with knowledge of different programming languages;

[0138] In this embodiment, the low-resource programming language model construction module is used to construct a low-resource programming language model, including an input layer, an encoding layer, a pooling layer and a classification layer. The input layer is used to input instances in the code corpus, the encoding layer calculates the semantic representation of the instances in the code corpus, the encoding layer is initialized by the pre-trained Adapter network corresponding to the programming language knowledge, the pooling layer calculates the semantic vector representation of the instance, and the classification layer calculates the prediction result for the masked words;

[0139] In this embodiment, the low-resource programming language model training module is used to obtain the instance x in the code corpus. c , respectively construct corresponding positive and negative examples, and use classification cost and contrastive learning cost to jointly train the low-resource programming language model;

[0140] In this embodiment, the code pre-training result output module is used to output the code pre-training result based on the final code pre-training model.

[0141] Example 3

[0142] This embodiment provides a storage medium, which may be a storage medium such as a ROM, RAM, a disk, or an optical disk. The storage medium stores one or more programs. When the program is executed by a processor, the code pre-training method based on an adapter network and contrastive learning of Embodiment 1 is implemented.

[0143] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principles of the present invention shall be equivalent replacement methods and shall be included in the protection scope of the present invention.

Claims

1. A code pre-training method based on Adapter network and contrastive learning, characterized in that: The steps include: Build a code pre-training model based on the Adapter network, including: multiple Transformer layers, Adapter networks corresponding to different programming languages, fusion layers and classification layers. Multiple Transformer layers learn general knowledge of programming languages, the Adapter network learns the domain knowledge of the current programming language, the fusion layer integrates general knowledge and domain knowledge of different programming languages, and the classification layer calculates the prediction results of the mask language; Get the training instance x in the code corpus d , train the code pre-training model based on the Adapter network to obtain the Adapter network with knowledge of different programming languages; Construct a low-resource programming language model, including an input layer, an encoding layer, a pooling layer, and a classification layer. The input layer inputs instances in the code corpus, and the encoding layer calculates the semantic representation of instances in the code corpus. The encoding layer is initialized by the pre-trained Adapter network corresponding to the programming language knowledge. The pooling layer calculates the semantic vector representation of the instance, and the classification layer calculates the prediction results for the masked words. Get instance x in the code corpus c , respectively construct corresponding positive and negative examples, and use classification cost and contrastive learning cost to jointly train the low-resource programming language model; The final code pre-training model is obtained through training, and the code pre-training result is output based on the final code pre-training model.

2. The code pre-training method based on adapter network and contrastive learning according to claim 1, characterized in that: The general knowledge of the programming language includes: data structures and algorithms, as well as programming paradigms and control flow structures.

3. The code pre-training method based on adapter network and contrastive learning according to claim 1, characterized in that: Get the training instance x in the code corpus d , training the code pre-training model based on the Adapter network, including: The Transformer layer is initialized based on the parameters of the code pre-training model CodeBERT; Input training instance x d , after multiple Transformer layers, the corresponding output is obtained, which is specifically expressed as: in, and They represent the kth and 12th Transformer layers of the code pre-training model CodeBERT, respectively. and Represents the outputs corresponding to the kth and 12th Transformer layers of the code pre-training model CodeBERT; Each layer of the Adapter network consists of a multi-layer feedforward neural network and a Transformer layer; Given a programming language p and an instance x d The initial semantic matrix representation of And the output of the first Transformer layer of the code pre-trained model CodeBERT The Adapter network is transformed as follows: ... Among them, FFN1, FFN2 and Represents the 1st, 2nd and Nth p Multi-layer feed-forward neural network in the Adapter layer, The first, second and Nth p The Transformer layer in the Adapter layer, and The first, second and Nth p The output of the Adapter layer, [] represents the matrix concatenation operation; The fusion layer includes a concatenation layer and an MLP layer. The fusion layer calculates its final semantic vector representation h as follows: d,p : Among them, [] represents the matrix concatenation operation of the concatenation layer, and MLP represents the MLP layer; The classification layer consists of a fully connected layer and a softmax transformation to obtain the instance x in the code corpus of a given programming language p d The final semantic vector representation h d,p After that, the classification layer calculates the prediction result of the mask language Specifically expressed as: Among them, W1 and b1 are the parameters in the fully connected layer, It is a v-dimensional vector, where v represents the size of the vocabulary during pre-training.

4. The code pre-training method based on adapter network and contrastive learning according to claim 3, characterized in that: When training the code pre-training model based on the Adapter network, given the code corpus D c The example in (x d ,y d ), the cross entropy cost function of the masked language model is defined as follows: Among them, γ c represents the parameter set of the entire model, Represents the expected value of the predicted result with respect to the true label.

5. The code pre-training method based on adapter network and contrastive learning according to claim 1, characterized in that: The encoding layer calculates the semantic representation of the instances in the code corpus, which is specifically expressed as: G c =Adapter p (x c ) Among them, H c , and For instance x c , positive example and negative examples The semantic matrix representation of , each column in the matrix corresponds to a word in the instance, Adapter p represents the coding layer; The pooling layer calculates the semantic vector representation of the instance, specifically: h c =AvgPooling(H c ) Among them, h c , and For instance x c , positive example and negative examples The semantic vector representation of , AvgPooling is the average pooling operation; The classification layer includes a fully connected layer and a softmax transformation. c The semantic vector representation h c After that, the classification layer of the masked language model calculates the prediction result for the masked word Specifically expressed as: Among them, W2 and b2 are the parameters of the fully connected layer, is a v-dimensional vector, where v is the size of the vocabulary in the corpus.

6. The code pre-training method based on adapter network and contrastive learning according to claim 5, characterized in that: The classification cost adopts the cross entropy cost function, which is specifically expressed as: Among them, θ′ c represents the parameter set of the connective classification model, y c For instance x c The true category to which the masked word belongs, It represents the expected value of the predicted word category with respect to the category of the real word.

7. The code pre-training method based on adapter network and contrastive learning according to claim 5, characterized in that: The contrastive learning cost is expressed as: Among them, h c , and For instance x c , positive example and negative examples The semantic vector representation of , sim() is used to measure the cosine distance between two vectors, |||| represents the 2-norm of the vector, and T represents the transpose of the vector.

8. The code pre-training method based on adapter network and contrastive learning according to claim 5, characterized in that: The total cost of the model is obtained by linearly summing the classification cost and the contrastive learning cost, which is expressed as: L(θ′ c )=L MLM (θ′ c )+λL cl (θ′ c ) Among them, λ is the weight coefficient, which is used to adjust the importance of pre-training cost and comparative learning cost, L MLM (θ′ c ) represents the classification cost, L cl (θ′ c ) represents the contrastive learning cost.

9. A code pre-training system based on Adapter network and contrastive learning, characterized in that: include: Adapter network-based code pre-training model construction module, Adapter network-based code pre-training model training module, low-resource programming language model construction module, low-resource programming language model training module, code pre-training result output module; The Adapter network-based code pre-training model construction module is used to construct a code pre-training model based on the Adapter network, including: multiple Transformer layers, Adapter networks corresponding to different programming languages, a fusion layer and a classification layer, the multiple Transformer layers learn general knowledge of programming languages, the Adapter network learns the domain knowledge of the current programming language, the fusion layer fuses general knowledge and domain knowledge of different programming languages, and the classification layer calculates the prediction result of the mask language; The code pre-training model training module based on the Adapter network is used to obtain the training instance x in the code corpus. d , train the code pre-training model based on the Adapter network to obtain the Adapter network with knowledge of different programming languages; The low-resource programming language model construction module is used to construct a low-resource programming language model, including an input layer, an encoding layer, a pooling layer and a classification layer, the input layer inputs instances in the code corpus, the encoding layer calculates semantic representations of instances in the code corpus, the encoding layer is initialized by a pre-trained Adapter network corresponding to programming language knowledge, the pooling layer calculates semantic vector representations of instances, and the classification layer calculates prediction results for masked words; The low-resource programming language model training module is used to obtain instances x in the code corpus. c , respectively construct corresponding positive and negative examples, and use classification cost and contrastive learning cost to jointly train the low-resource programming language model; The code pre-training result output module is used to output the code pre-training result based on the final code pre-training model.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the code pre-training method based on an adapter network and contrastive learning as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-language translation model construction method and storage medium

    CN115688815A

  • Implicit discourse relation identification method and system based on comparative learning and Adapter network

    CN116028630A

  • Inter-training of pre-trained transformer-based language models using partitioning and classification

    US20230078698A1