Code generation method, device, storage medium and electronic device

By integrating code understanding tasks and code generation tasks in the code generation model, the problem of difficult to capture code semantics in the existing technology is solved, and more accurate and high-quality code generation is achieved.

CN116166271BActive Publication Date: 2025-05-16DOUYIN VISION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310190548.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-05-16
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively learn and capture the semantic information of the code in program code generation, and can only learn the statistical distribution of symbols, and cannot effectively capture the code structure information.

Method used

The multi-task learning method is adopted to integrate code understanding tasks and code generation tasks in the training process of the code generation model, and learn the syntax and semantic characteristics of sample program code through the code understanding task, and then generate new program code in the code generation task.

Benefits of technology

By combining syntax and semantic knowledge to generate code, the accuracy and quality of code generation is improved, and the syntax and semantic features of the target text can be captured more effectively to generate more accurate program code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166271B_ABST
    Figure CN116166271B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a code generation method, device, storage medium and electronic device to improve the accuracy of automatically generated code. The method comprises: obtaining a target text, the target text comprising a program code text to be supplemented or a natural language text for describing the function of the code; inputting the target text into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model is trained by a code understanding task and a code generation task, and the code understanding task is used by the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used by the code generation model to learn the process of generating a new program code based on the sample program code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a code generation method, device, storage medium and electronic device. Background Art

[0002] Automatic program code generation is an important technical means for software intelligence. It can assist software developers in completing the development of some general codes, thereby effectively improving R&D efficiency and saving R&D costs. In addition, with the continuous development of natural language processing technology, pre-trained language models are widely used in program code generation. However, related technologies mainly generate code based on conditional probabilistic language models, that is, predicting the probability of the next program symbol based on the existing context information, and can only learn the statistical distribution of program symbols. Summary of the invention

[0003] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a code generation method, the method comprising:

[0005] Acquire a target text, wherein the target text includes a program code text to be supplemented or a natural language text for describing a code function;

[0006] The target text is input into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model is trained by a code comprehension task and a code generation task, and the code comprehension task is used for the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used for the code generation model to learn the process of generating a new program code based on the sample program code.

[0007] In a second aspect, the present disclosure provides a code generation device, the device comprising:

[0008] An acquisition module, used for acquiring a target text, wherein the target text is a program code text to be supplemented or a natural language text for describing a code function;

[0009] A generation module is used to input the target text into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model is trained by a code comprehension task and a code generation task, and the code comprehension task is used for the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used for the code generation model to learn the process of generating a new program code based on the sample program code.

[0010] In a third aspect, the present disclosure provides a non-transitory computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processing device.

[0011] In a fourth aspect, the present disclosure provides an electronic device, including:

[0012] a storage device having a computer program stored thereon;

[0013] A processing device is used to execute the computer program in the storage device to implement the steps of the method in the first aspect.

[0014] Through the above technical solution, the code generation model can be trained through code understanding tasks and code generation tasks. Among them, the code understanding task is used for the code generation model to learn the grammar and semantic knowledge of the sample program code, and the code generation task is used for the code generation model to learn the process of generating new program code based on the sample program code. Therefore, the trained code generation model can generate the target program code according to the grammar and semantics of the target text, thereby improving the accuracy of the automatic code generation.

[0015] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings:

[0017] Figure 1 is a flowchart of a code generation method according to an exemplary embodiment of the present disclosure;

[0018] Figure 2 is a process schematic diagram of a code generation method according to an exemplary embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram of a code generation model in a code generation method according to an exemplary embodiment of the present disclosure;

[0020] Figure 4 is a process schematic diagram of a code generation method according to another exemplary embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of a mask matrix in a code generation method according to an exemplary embodiment of the present disclosure;

[0022] Figure 6 is a process schematic diagram of a code generation method according to another exemplary embodiment of the present disclosure;

[0023] Figure 7 is a block diagram of a code generating device according to an exemplary embodiment of the present disclosure;

[0024] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0026] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0027] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0028] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should also be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0030] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0031] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0032] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0033] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0034] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0035] As mentioned in the background technology, with the continuous development of natural language processing technology, pre-trained language models are widely used in program code generation. However, this type of method regards program code as a symbol sequence similar to natural language text, and uses conditional probability autoregressive generation to train code generation. It cannot effectively learn the semantics of program code, but can only learn the statistical distribution of program symbols, and the structural information of program code cannot be effectively captured.

[0036] Based on this, the present disclosure provides a code generation method that can integrate a multi-task learning method of program code understanding and program code generation in the training process of the code generation model. That is to say, in order to enable the model to effectively learn the grammatical and semantic knowledge of the program code, the present disclosure sets two major types of training tasks in terms of training objectives: code understanding tasks and code generation tasks. In the code understanding task, the model can learn the basic grammatical knowledge and semantic knowledge of the program code, so that in the model application stage, the model can generate program code according to the grammar and semantics of the input text, thereby improving the accuracy of the generated code.

[0037] Figure 1 is a flowchart of a code generation method according to an exemplary embodiment of the present disclosure. Figure 1 , the code generation method may include:

[0038] Step 101 : obtaining a target text, where the target text includes a program code text to be supplemented or a natural language text for describing a code function.

[0039] Step 102, input the target text into the code generation model to obtain the target program code generated based on the target text. The code generation model is trained by the code understanding task and the code generation task, and the code understanding task is used for the code generation model to learn the grammatical features and semantic features of the sample program code, and the code generation task is used for the code generation model to learn the process of generating new program code based on the sample program code.

[0040] It should be understood that the target text input into the code generation model in the present disclosure may include a program code to be supplemented or a natural language text for describing the code function. Among them, the program code to be supplemented may be an incomplete program code. After the program code to be supplemented is input into the code generation model, the code generation model can output a complete program code according to the program code to be supplemented. The natural language text used to describe the code function, for example, can be a natural language text determined according to the actual business function. After the natural language text is input into the code generation model, the code generation model can output a program code that can realize the business function. Thus, the automatic generation of code can be realized, the efficiency of code generation can be improved, and the code understanding and code generation pre-training tasks are introduced simultaneously in the training stage of the code generation model, so that the code generation model can effectively capture the grammatical and semantic knowledge in the target text in the application stage to enhance the performance of code generation and improve the accuracy of the generated program code.

[0041] The following first describes the training process of the code generation model.

[0042] First, a sample program code may be obtained, then the sample program code is tokenized and split into a fixed-length token sequence, and finally the code generation model is trained based on the token sequence, wherein the fixed length may be 1024 tokens (tokens).

[0043] In some embodiments, the sample program code can be obtained in the following manner: first obtain a code data set, which includes multiple code files. Then, for each code file, perform at least one of the following preprocessing steps to obtain a training data set: when the ratio of the number of characters in the code file to the total number of symbols is greater than or equal to a preset character ratio, add the code file to the training data set; when the average number of characters per line of code in the code file is less than or equal to a first preset threshold, add the code file to the training data set; when the number of characters in the comment information in the code file is less than or equal to a second preset threshold, add the code file to the training data set. Finally, the program code corresponding to each code file in the training data set is determined as the sample program code.

[0044] The preset character ratio, the first preset threshold and the second preset threshold can be set according to actual conditions, and the embodiments of the present disclosure do not limit this. For example, the preset character ratio can be set to 40%, the first preset threshold can be set to 100, and the second preset threshold can be set to 500.

[0045] It should be understood that the directly acquired code data set usually contains more repeated code data and low-quality code data. In order to prevent these repeated code data and low-quality data from affecting the learning process of the model, the acquired code data set can be preprocessed first.

[0046] For example, refer to Figure 2First, traverse the code files in the code data set in turn and calculate the MD5 (Message-Digest Algorithm) hash value of the file. After the hash value is calculated for the first time, it is stored. After the hash value is calculated for the second time, it is compared with the stored hash value. If the hash value exists, it indicates that it is a duplicate file and can be discarded. If the hash value does not exist, the hash value of this time is stored. And so on, the code files in the code data set are deduplicated. After that, the ratio of the number of characters in the file (the number of characters consisting of at least one letter of the 26 English letters) to the total number of symbols in the code file (the number of letters, spaces, and other special characters, one space is counted as one character) and the average number of characters per line in the code file can be calculated. If the ratio of the number of characters in the code file to the total number of symbols is less than 40% or the average number of characters per line in the code file exceeds 100, it can be considered that this is a code file containing more noise, so the code file is not added to the training data set, and the next code file is calculated. On the contrary, if the number of characters in the code file accounts for a proportion of the total number of symbols greater than or equal to 40% or the average number of characters per line in the code file does not exceed 100, the code file can be added to the training data set. Then, in the preprocessing stage, the annotation information of more than 500 characters in the code file can be removed to improve the quality of the training data. Finally, the program code corresponding to each code file in the training data set is determined as the sample program code.

[0047] Through the above method, we can filter out duplicate and low-quality code data based on the original code data and establish a high-quality training data set, thereby reducing the impact of duplicate and low-quality code data on model training, improving the training effect, and further improving the accuracy of the results of the trained code generation model.

[0048] In some embodiments, the code generation model includes a code understanding layer, a code generation layer and a shared representation layer, and the result output by the shared representation layer can be used to input the code understanding layer or the code generation layer. Accordingly, the training step of the code generation model includes: when performing a code understanding task, masking the sample program code to obtain a mask sequence, inputting the mask sequence into the code generation model to obtain a first prediction code through the shared representation layer and the code understanding layer, and adjusting the parameters of the code understanding layer and the shared representation layer based on the first prediction code. When performing a code generation task, inputting the sample program code into the code generation model to obtain a second prediction code through the shared representation layer and the code generation layer, and adjusting the parameters of the code generation layer and the shared representation layer based on the result output by the code generation model.

[0049] That is to say, when performing code comprehension tasks, training can be performed using a masked language model, and when performing code generation tasks, training can be performed using a conditional language model.

[0050] For example, refer to Figure 3 The code generation model includes a shared representation layer, a code understanding layer, and a code generation layer. In the training phase, the parameters of the shared representation layer and the code understanding layer can be adjusted through the code understanding task, and then the parameters of the shared representation layer and the code generation layer can be adjusted through the code generation task. Therefore, in the model application phase, the grammatical and semantic knowledge of the target text can be identified, so that more accurate target program code can be output through the grammatical and semantic knowledge.

[0051] The following describes the training process of the code comprehension task.

[0052] In some embodiments, masking is performed on the sample program code to obtain a masked sequence, including: performing word segmentation processing on the sample program code to obtain a word segmentation sequence; randomly selecting a first preset proportion of first target word segments in the word segmentation sequence, and replacing the first target word segments with a first preset mask symbol to obtain a masked sequence; and / or randomly selecting a second preset proportion of second target word segments in the word segmentation sequence, and for each second target word segmentation, randomly selecting a target remaining word segmentation from the remaining word segments, replacing the second target word segmentation with the target remaining word segmentation to obtain a masked sequence.

[0053] Among them, the remaining segmentation is the segmentation in the segmentation sequence except the second target segmentation. The first preset ratio, the second preset ratio and the first preset mask symbol can be set according to the actual situation, and the embodiment of the present disclosure does not limit this. For example, the first preset ratio is set to 12%, the second preset ratio is set to 1.5%, and the first preset mask symbol is set to [MASK]. In addition, the segmentation of the first preset ratio and the second preset ratio can be directly selected, or a certain proportion of segmentation can be selected first, and then a certain proportion of segmentation can be selected from the selected segmentation. For example, after determining the segmentation sequence corresponding to the training sample, 15% of the tokens in the segmentation sequence are randomly selected in units of tokens, and then 80% of the tokens (i.e., the first target segmentation) in the selected tokens are replaced with the [MASK] symbol, and 10% of the tokens (i.e., the second target segmentation) in the selected tokens are randomly replaced with other tokens to obtain a mask sequence. Therefore, by training the code generation model with the mask sequence, the model can learn to recover the information of the mask segmentation from the mask sequence, so that the model can learn the basic grammatical knowledge of the program code.

[0054] In some embodiments, masking is performed on the sample program code to obtain a mask sequence, including: performing word segmentation on the sample program code to obtain a word segmentation sequence. Then, function word segmentations used to characterize function names are determined in the word segmentation sequence, and a third preset ratio of target function word segmentations are randomly selected from the function word segmentations, and the target function word segmentations are replaced with a second preset mask symbol to obtain a mask sequence; and / or, interface word segmentations used to characterize interface names are determined in the word segmentation sequence, and a fourth preset ratio of target interface word segmentations are randomly selected from the interface word segmentations, and the target interface word segmentations are replaced with a third preset mask symbol to obtain a mask sequence.

[0055] Among them, the third preset ratio, the fourth preset ratio, the second preset mask symbol and the third preset mask symbol can be set according to actual conditions, and the embodiments of the present disclosure do not limit this. For example, the third preset ratio and the fourth preset ratio are both set to 20%, and the third preset mask symbol and the fourth preset mask symbol are both set to [MASK], that is, the token of the function name in the sample program code can be replaced with the [MASK] symbol at a ratio of 20%, and / or, the token of the interface (API, Application Programming Interface) name in the sample program code can be replaced with the [MASK] symbol at a ratio of 20% to obtain a mask sequence. Afterwards, the code generation model is trained with the mask sequence, so that the model can learn to restore the masked function name and API name, so that the model can learn the correspondence between the function name and the function body and semantic knowledge such as the application program interface.

[0056] In practical applications, mask processing can be performed by any of the above methods, or by combining the above two methods, which is not limited in the embodiments of the present disclosure.

[0057] After obtaining the mask sequence, the mask sequence can be input into the code generation model to obtain the first predicted code through the shared representation layer and the code understanding layer. In some embodiments, the feature vector sequence corresponding to the mask sequence can be determined first. Then, for each feature vector in the feature vector sequence, the shared representation layer is used to calculate all feature vectors before and after the feature vector in the feature vector sequence and the attention mechanism to obtain an intermediate feature vector, and then for each intermediate feature vector, the code understanding layer is used to perform attention calculation based on the intermediate feature vectors other than the intermediate feature vector and the attention mechanism to obtain a target feature vector. Finally, the first predicted program code is obtained based on the target feature vector.

[0058] Among them, the shared representation layer performs calculations based on all feature vectors before and after the feature vector in the feature vector sequence and the attention mechanism, which is equivalent to the shared representation layer performing a two-way attention calculation. The code understanding layer performs attention calculations based on the intermediate feature vectors other than the intermediate feature vectors and the attention mechanism, which is equivalent to the code understanding layer performing a two-way attention calculation. Therefore, when performing the code understanding task, both the shared representation layer and the code understanding layer perform a two-way attention calculation, so that the code understanding can be combined with the previous and following contexts to more accurately identify the grammatical and semantic knowledge of the input text.

[0059] For example, refer to Figure 4 , taking the code generation model as Transformer as an example to illustrate the execution process of each layer of the model in the training phase. First, the code snippet after the sample program code is segmented passes through the word embedding layer and the position encoding layer to obtain the word vector representation of each word. Among them, the word embedding layer obtains the corresponding word embedding vector from the word embedding matrix by looking up the table. The word vector matrix is ​​a matrix of size 42000×1024, where 42000 is the size of the vocabulary and 1024 is the dimension of the word vector. The position encoding layer calculates the position encoding vector of each word through the cosine-based position encoding method. Finally, the word embedding vector and the position encoding vector corresponding to each word are added to obtain the word vector of the word, as shown in formula (1):

[0060]

[0061] in, Represents the word vector corresponding to the i-th word, WordEmbedding(i) represents the word embedding vector corresponding to the i-th word, and PositionEmbedding(i) represents the position encoding vector corresponding to the i-th word.

[0062] The sequence after the word embedding layer is input into the shared representation layer for encoding. The shared representation layer can be composed of multiple transformer basic blocks (transformerblock) stacked together, each of which consists of a multi-head attention sublayer and a forward feedback sublayer ( Figure 4 The multi-head attention sublayer is used to calculate the correlation between each word in the input sequence. In practical applications, the number of attention heads can be set to 16, that is, 16 attention heads can be used, which is not limited in the embodiment of the present disclosure. In each attention head, the encoding representation h of the i-th word can be calculated by the self-attention mechanism i , as shown in formula (2):

[0063]

[0064] Among them, Q represents the word vector H of the i-th word iThe query vector obtained after embedding is linearly transformed, K represents the word vector H of the i-th word i The key vector obtained after the linear transformation of embedding, V represents the word vector H of the i-th word i The value vector obtained after the linear transformation of embedding, K T represents the transpose of K, d k Represents the dimension of K.

[0065] After obtaining the encoding representation h of a single attention head output i After that, the outputs of multiple attention heads are concatenated and passed through a fully connected layer to obtain the encoding representation m of the i-th word i , as shown in formula (3):

[0066] m i =concat(h′1,..,h′ h )W o (3)

[0067] Among them, concat represents vector concatenation, h′1 represents the encoding representation of the output of the first attention head, and h′ h represents the encoded representation of the h-th attention head output, W o Represents the parameters of the multi-head attention sub-layer.

[0068] Furthermore, to prevent the gradient from vanishing, the output of the multi-head attention sublayer can be obtained after a residual connection and a normalization layer. As shown in formula (4):

[0069]

[0070] Among them, LayerNorm represents the normalization function.

[0071] Afterwards, the forward feedback sublayer further processes the vector output by the multi-head attention sublayer. The forward feedback sublayer consists of two fully connected units and one ReLU activation unit, and its calculation process is shown in formula (5):

[0072]

[0073] Among them, max represents the maximum value function, h″ i represents the output of the forward feedback sublayer, W1 and W2 represent the parameters of the forward feedback sublayer, and b1 and b2 are the corresponding biases. It should be understood that the specific calculation process can refer to the calculation process of the transformer model in the relevant technology, which will not be repeated here.

[0074] Similarly, the output of the forward feedback sublayer also passes through the residual connection and normalization layer to obtain the output of the shared representation layer. As shown in formula (6):

[0075]

[0076] Therefore, the shared representation layer can perform bidirectional attention calculation based on all feature vectors before and after the feature vector in the feature vector sequence to obtain the intermediate feature vector, that is,

[0077] After that, for each intermediate feature vector, the code understanding layer can perform a bidirectional attention calculation based on the intermediate feature vectors other than the intermediate feature vector to obtain the target feature vector. For example, using the above example, the output of the shared representation layer is After the bidirectional attention calculation of the code understanding sub-network, the corresponding target feature vector is obtained after passing through the residual layer and the forward feedback layer in sequence. The calculation process is similar to that of the shared representation layer. Please refer to the previous article and will not be repeated here.

[0078] Finally, a first predicted program code is obtained according to the target feature vector, so that parameters of the code understanding layer and the shared representation layer are adjusted based on the first predicted program code.

[0079] For example, following the above example, the masked sequence after masking is input into the code generation model to obtain the feature representation h of each word. s . Then the corresponding words of the masked words are After the first linear transformation layer, the probability of the word belonging to different words in the vocabulary is obtained Then the probability is calculated through the softmax function to obtain the normalized probability p u , as shown in equations (7) and (8):

[0080]

[0081]

[0082] Among them, W u represents the parameters of the first linear transformation layer, b u represents the corresponding bias, m represents the masked word, and m i represents the label of the correct word corresponding to the masked word, θ represents the parameters of the code generation model, and N represents the total number of words in the vocabulary.

[0083] Then, the loss function of the code comprehension task is calculated as shown in formula (9):

[0084]

[0085] Among them, L1 represents the loss function value calculated in the code comprehension task, and M represents the number of words processed by masking.

[0086] Finally, the parameters of the code understanding layer and the shared representation layer can be adjusted according to the L1 loss function value. This process is similar to the related technology and will not be repeated here.

[0087] Therefore, when performing code comprehension tasks, both the shared representation layer and the code comprehension layer perform bidirectional attention calculations, so that the code can be understood by combining the previous and following information, and the grammatical and semantic knowledge of the input text can be more accurately identified.

[0088] The following describes the training process of the code generation task.

[0089] When executing the code generation task, the sample program code can be input into the code generation model to obtain the second prediction code through the shared representation layer and the code generation layer, and the parameters of the code generation layer and the shared representation layer are adjusted based on the results output by the code generation model.

[0090] In some embodiments, the sample program code is input into the code generation model to obtain the second prediction code through the shared representation layer and the code generation layer, including: first determining the feature vector sequence corresponding to the sample program code, wherein the feature vector sequence includes N feature vectors, N is a positive integer. Then, for the i-th feature vector in the feature vector sequence, the shared representation layer is used to calculate based on all feature vectors before and after the i-th feature vector in the feature vector sequence and the attention mechanism to obtain the i-th intermediate feature vector, i is a positive integer. For the i-th intermediate feature vector, the code generation layer is used to calculate based on the preceding feature vector and the attention mechanism to obtain the target feature vector, wherein the preceding feature vector is all the intermediate feature vectors obtained before the i-th intermediate feature vector. Finally, according to the target feature vector, the second prediction program code is obtained.

[0091] Among them, the shared representation layer performs calculations based on all feature vectors and attention mechanisms before and after the i-th feature vector in the feature vector sequence, which is equivalent to the shared representation layer performing bidirectional attention calculations. The code generation layer performs calculations based on the preceding feature vector and attention mechanisms, which is equivalent to the code generation layer performing unidirectional attention calculations. In other words, when executing the code generation task, the shared representation layer performs bidirectional attention calculations, and the code understanding layer performs unidirectional attention calculations, so that the next word can be predicted based on the previous context information of the word to achieve code generation. In addition, since the shared representation layer can adjust parameters through the code understanding task, the shared representation layer can combine the grammatical and semantic knowledge of the previous context information to obtain the corresponding features and input them into the code generation layer for processing, and then can combine the grammatical and semantic knowledge of the input text to perform code generation, thereby improving the accuracy of code generation.

[0092] For example, refer to Figure 3 The multi-task layer includes two sub-networks, code understanding and code generation, and the two sub-networks are composed of multiple transformer basic blocks. The difference is that the code understanding sub-network uses the same bidirectional attention calculation method as the shared representation layer, that is, when calculating the dependency of a word, it will calculate the dependency between it and the previous and next words, while the code generation sub-network uses a unidirectional attention calculation method, that is, when calculating the dependency of a word, it will only calculate the dependency between the word and the previous word. For example, the code generation sub-network uses a mask-based attention mechanism to calculate the feature representation of words, as shown in formula (10):

[0093]

[0094] Among them, Q s express The query vector obtained after linear transformation, K s express The key vector obtained after linear transformation, V s express The value vector obtained after linear change, K T represents the transpose of K, d′ k K s The dimension, M ′ represents a mask matrix of size n×n, where n represents the length of the input sequence, and the values ​​of all elements on the diagonal and to the left of the matrix are 1, while all elements on the right of the diagonal are 0. For example, Figure 5 is an example of a mask matrix for an input sequence of length 6.

[0095] It should be understood that the other calculation processes of the code generation task subnetwork can refer to the calculation process of the native transformer in the relevant technology, which will not be repeated here.

[0096] As mentioned above, in the code generation subnetwork, the mask matrix is ​​introduced so that the feature output of the i-th word is It is obtained by weighted summation of its predecessor feature vectors, so it represents the feature representation of the predecessor sequence. Therefore, the feature output of each word calculated in the code generation subnetwork can be After the second linear transformation layer, the probability of each word belonging to a different word in the vocabulary is obtained After the softmax function is used to calculate, the normalized probability P of each word in the vocabulary is obtained. g , as shown in equations (11) and (12)

[0097]

[0098]

[0099] Among them, W g represents the parameters of the second linear transformation layer, b g represents the corresponding bias, m′ represents the next predicted word to be generated, and m′ i Indicates the label information of the correct word corresponding to the predicted word.

[0100] Then, the loss function of the code generation task is calculated as shown in formula (13):

[0101]

[0102] Wherein, L2 represents the loss function value calculated in the code generation task, and M″ represents the number of predicted words generated.

[0103] Finally, the parameters of the code generation layer and the shared representation layer can be adjusted according to the L2 loss function value. This process is similar to the related technology and will not be repeated here.

[0104] In the above manner, during the training process of the code generation model, the code understanding task and the code generation task can be iteratively executed. For example, in the first round of iterations, the code understanding task is first executed to train the model through the L1 loss function value, and only the parameters of the shared representation layer and the code understanding layer of the model are updated during this process. Subsequently, in the second round of iterations, the code generation task is executed to train the model through the L2 loss function value, and only the parameters of the shared representation layer and the code generation layer are updated during this process. This cycle is repeated until the conditions for the end of model training are met, such as when both the L1 loss function value and the L2 loss function value are less than a preset threshold. In addition, during the entire training process, the gradient descent method can be used to update the corresponding parameters of the model, which is not limited in the embodiments of the present disclosure.

[0105] The following describes the application process of the code generation model.

[0106] After the code generation model is trained in the above manner, the natural language sentences or the program code fragments to be supplemented can be segmented to obtain a sequence of segmented words to be processed, and then the sequence of segmented words to be processed is input into the trained code generation model to obtain a new target program code.

[0107] It should be understood that in the application stage, the code generation model only uses the shared representation layer and the code generation layer to predict the probability of each word in the target program code. The steps may include: 1. Segmenting the natural language sentence or the program code fragment to be supplemented to obtain the segmented word sequence to be processed. 2. Inputting the segmented word sequence to be processed into the code generation model, and obtaining the feature representation of the next predicted word after encoding. 3. Represent the features of the next predicted word After being sent to the linear decoding layer, the probability distribution P of the predicted word in the vocabulary is obtained after the softmax function is calculated. g ′, the specific calculation process can refer to formulas (11) and (12). 4. After obtaining the probability distribution P g ', select the target predicted word segmentation to be output according to the probability. 5. Connect the target predicted word segmentation to the word segmentation sequence to be processed and input it into the model. Repeat steps 3 and 4 to obtain a new predicted word segmentation. Stop generating when the preset maximum sequence length is reached or the output predicted word segmentation is the end symbol.

[0108] In some embodiments, the segmented word with the highest probability may be selected as the target predicted segmented word for output in step 4. Alternatively, in some embodiments, in order to increase the diversity of the output target predicted segmented words, the predicted word may be obtained by a nucleus sampling decoding algorithm.

[0109] For example, the code generation model can be used to obtain the target program code generated based on the target text in the following manner: first determine a predicted word based on the target text, and then use the predicted word as the initial target word to loop through the following process: concatenate the target word in the word segmentation sequence corresponding to the target text to obtain the target word segmentation sequence, determine multiple candidate predicted words and the probability of each candidate predicted word in the preset word list based on the target word segmentation sequence, and randomly select a word from the candidate set corresponding to the candidate predicted word as the new target word, until the length of the target word segmentation sequence reaches the preset length or the target word is the preset end symbol. Among them, the sum of the probabilities of the candidate predicted words in the candidate set is greater than the preset probability.

[0110] For example, when we get the probability distribution P g ′, first from the probability distribution P g′, select a candidate set V p , so that the sum of the probabilities of all candidate predicted words in the candidate set is greater than the preset probability p, and then the candidate set V p The probabilities of all candidate predicted words in are normalized, and a word is randomly sampled as the target word. The calculation process is shown in formula (14):

[0111]

[0112] Wherein, x represents the predicted word, and the preset probability p can be set according to the actual situation, for example, it can be set to 0.95, which is not limited in the embodiment of the present disclosure.

[0113] After a target word is predicted, the target word is concatenated with the word segmentation sequence input into the code generation model to obtain a target word segmentation sequence, which is then input into the code generation model, and steps (3) and (4) are repeated to predict and output a new target word until the length of the target word segmentation sequence reaches a preset length or the output target word is a preset end symbol. The preset length and the preset end symbol can be set according to actual conditions, and the embodiments of the present disclosure do not limit this.

[0114] For example, refer to Figure 6 , the code generation method provided by the present disclosure mainly includes three processes: data preprocessing, model training and model application. It should be understood that the specific processing methods of the three processes can refer to the above, and are briefly described below. First, in the data preprocessing process, a program code data set can be obtained, and then the program code data set is subjected to data denoising, and a vocabulary is trained through a BPE (Byte Pair Encoding) word segmentation algorithm to obtain a vocabulary for subsequent word segmentation. Afterwards, the denoised data is word segmented and split into sequence fragments of fixed length through the trained vocabulary, and the sequence fragments of fixed length are input into the code generation model for training. In the model training process, the code generation model is trained based on the code understanding task and the code generation task. In the model application process, for the program code fragment to be supplemented, a word segmentation process is first performed to obtain a word segmentation sequence to be processed, and then the word segmentation sequence to be processed is input into the trained code generation model. Afterwards, the final target program code is obtained by a nuclear sampling decoding algorithm.

[0115] Therefore, the code generation model can be trained based on a multi-task learning method of code understanding and code generation. Among them, the code understanding task can use a masked language model to let the model learn to recover word information from masked text, and the masked words are randomly selected, so that the model can learn the basic grammatical knowledge of the program code. On the other hand, the code understanding task can allow the model to learn to recover the masked function names and API names, so as to learn semantic knowledge such as the correspondence between the function name and the function body. In the code generation task, an autoregressive conditional language model can be used for training, that is, the next most likely code text symbol is predicted through the above information. Therefore, the trained code generation model can output more accurate target program code based on the program code text to be supplemented or the natural language text used to describe the code function.

[0116] Based on the same concept, the present disclosure also provides a code generation device, which can be part or all of an electronic device through software, hardware, or a combination of both. Figure 7 , the code generating device 700 comprises:

[0117] An acquisition module 701 is used to acquire a target text, where the target text is a program code text to be supplemented or a natural language text used to describe a code function;

[0118] A generation module 702 is used to input the target text into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model is trained by a code comprehension task and a code generation task, and the code comprehension task is used for the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used for the code generation model to learn the process of generating a new program code based on the sample program code.

[0119] Optionally, the code generation model includes a code understanding layer, a code generation layer and a shared representation layer, and a result output by the shared representation layer is used to input the code understanding layer or the code generation layer. The apparatus 700 further includes a training module for:

[0120] When executing the code understanding task, masking is performed on the sample program code to obtain a mask sequence, the mask sequence is input into the code generation model to obtain a first prediction code through the shared representation layer and the code understanding layer, and parameters of the code understanding layer and the shared representation layer are adjusted based on the first prediction code;

[0121] When executing the code generation task, the sample program code is input into the code generation model to obtain a second predicted code through the shared representation layer and the code generation layer, and the parameters of the code generation layer and the shared representation layer are adjusted based on the result output by the code generation model.

[0122] Optionally, the training module is used to:

[0123] Performing word segmentation processing on the sample program code to obtain a word segmentation sequence;

[0124] Randomly selecting a first preset proportion of first target segmented words in the segmented word sequence, and replacing the first target segmented words with a first preset mask symbol to obtain a mask sequence; and / or,

[0125] A second preset proportion of second target participles is randomly selected from the participle sequence, and for each of the second target participles, a target remaining participle is randomly selected from the remaining participles, and the second target participles are replaced with the target remaining participles to obtain a mask sequence, wherein the remaining participles are the participles in the participle sequence except the second target participles.

[0126] Optionally, the training module is used to:

[0127] Performing word segmentation processing on the sample program code to obtain a word segmentation sequence;

[0128] Determining function participles used to represent function names in the participle sequence, and randomly selecting a third preset proportion of target function participles in the function participles, replacing the target function participles with second preset mask symbols, to obtain a mask sequence; and / or

[0129] An interface segmentation for representing an interface name is determined in the segmentation sequence, and a fourth preset proportion of target interface segmentations are randomly selected from the interface segmentations, and the target interface segmentations are replaced with a third preset mask symbol to obtain a mask sequence.

[0130] Optionally, the training module is used to:

[0131] Determine a feature vector sequence corresponding to the mask sequence;

[0132] For each feature vector in the feature vector sequence, the shared representation layer performs calculations based on all feature vectors before and after the feature vector in the feature vector sequence and the attention mechanism to obtain an intermediate feature vector;

[0133] For each of the intermediate feature vectors, the code understanding layer performs attention calculation according to the intermediate feature vectors other than the intermediate feature vector and the attention mechanism to obtain a target feature vector;

[0134] A first prediction program code is obtained according to the target feature vector.

[0135] Optionally, the training module is used to:

[0136] Determine a feature vector sequence corresponding to the sample program code, wherein the feature vector sequence includes N feature vectors, where N is a positive integer;

[0137] For the i-th feature vector in the feature vector sequence, the shared representation layer calculates according to all feature vectors before and after the i-th feature vector in the feature vector sequence and the attention mechanism to obtain an i-th intermediate feature vector, where i is a positive integer;

[0138] For the i-th intermediate feature vector, the code generation layer performs calculation according to the preceding feature vector and the attention mechanism to obtain a target feature vector, wherein the preceding feature vector is all the intermediate feature vectors obtained before the i-th intermediate feature vector;

[0139] A second prediction program code is obtained according to the target feature vector.

[0140] Optionally, the device 700 further includes a preprocessing module, configured to:

[0141] Acquire a code data set, wherein the code data set includes a plurality of code files;

[0142] For each of the code files, at least one of the following preprocessing steps is performed to obtain a target code data set: when the ratio of the number of characters in the code file to the total number of symbols is less than a preset character ratio, the code file is deleted from the code data set; when the average number of characters per line of code in the code file is greater than a first preset threshold, the code file is deleted from the code data set; when the number of characters in the comment information in the code file is greater than a second preset threshold, the code file is deleted from the code data set;

[0143] The program code corresponding to each of the code files in the target code data set is determined as a sample program code.

[0144] Optionally, the code generation model is used to obtain the target program code generated based on the target text in the following manner:

[0145] A predicted word is determined based on the target text, and the predicted word is used as the initial target word to loop through the following process:

[0146] The target word is concatenated into the word segmentation sequence corresponding to the target text to obtain a target word segmentation sequence, a plurality of candidate prediction words and the probability of each candidate prediction word in a preset word list are determined based on the target word segmentation sequence, and a word is randomly selected from the candidate set corresponding to the candidate prediction word as a new target word until the length of the target word segmentation sequence reaches a preset length or the target word is a preset end symbol;

[0147] Wherein, the sum of the probabilities of the candidate predicted words in the candidate set is greater than a preset probability.

[0148] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0149] Based on the same concept, the present disclosure also provides a non-transitory computer-readable medium on which a computer program is stored. When the program is executed by a processing device, the steps of any of the above code generation methods are implemented.

[0150] Based on the same concept, the present disclosure also provides an electronic device, including:

[0151] a storage device having a computer program stored thereon;

[0152] A processing device is used to execute the computer program in the storage device to implement the steps of any of the above code generation methods.

[0153] Reference below Figure 8 , which shows a schematic diagram of the structure of an electronic device 800 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0154] like Figure 8As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0155] Typically, the following devices may be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0156] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0157] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0158] In some embodiments, any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol) can be used for communication, and can be interconnected with any form or medium of digital data communication (e.g., communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), internets (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0159] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0160] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a target text, which includes a program code text to be supplemented or a natural language text used to describe a code function; inputs the target text into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model is trained by a code understanding task and a code generation task, and the code understanding task is used by the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used by the code generation model to learn the process of generating a new program code based on the sample program code.

[0161] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0162] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0163] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a module does not, in some cases, limit the module itself.

[0164] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0165] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0166] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0167] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0168] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.

Claims

1. A code generation method, characterized in that: The method comprises: Acquire a target text, wherein the target text includes a program code text to be supplemented or a natural language text for describing a code function; Inputting the target text into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model is trained by a code comprehension task and a code generation task, and the code comprehension task is used for the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used for the code generation model to learn a process of generating a new program code based on the sample program code; The code generation model includes a code understanding layer, a code generation layer and a shared representation layer. The result output by the shared representation layer is used to input the code understanding layer or the code generation layer. The training steps of the code generation model include: When executing the code understanding task, masking is performed on the sample program code to obtain a mask sequence, the mask sequence is input into the code generation model to obtain a first prediction code through the shared representation layer and the code understanding layer, and parameters of the code understanding layer and the shared representation layer are adjusted based on the first prediction code; When executing the code generation task, the sample program code is input into the code generation model to obtain a second predicted code through the shared representation layer and the code generation layer, and the parameters of the code generation layer and the shared representation layer are adjusted based on the result output by the code generation model.

2. The method according to claim 1, characterized in that The performing mask processing on the sample program code to obtain a mask sequence includes: Performing word segmentation processing on the sample program code to obtain a word segmentation sequence; Randomly selecting a first preset proportion of first target segmented words in the segmented word sequence, and replacing the first target segmented words with a first preset mask symbol to obtain a mask sequence; and / or, A second preset proportion of second target participles is randomly selected from the participle sequence, and for each of the second target participles, a target remaining participle is randomly selected from the remaining participles, and the second target participles are replaced with the target remaining participles to obtain a mask sequence, wherein the remaining participles are the participles in the participle sequence except the second target participles.

3. The method according to claim 1, characterized in that The performing mask processing on the sample program code to obtain a mask sequence includes: Performing word segmentation processing on the sample program code to obtain a word segmentation sequence; Determining function participles used to represent function names in the participle sequence, and randomly selecting a third preset proportion of target function participles in the function participles, replacing the target function participles with second preset mask symbols, to obtain a mask sequence; and / or An interface segmentation for representing an interface name is determined in the segmentation sequence, and a fourth preset proportion of target interface segmentations are randomly selected from the interface segmentations, and the target interface segmentations are replaced with a third preset mask symbol to obtain a mask sequence.

4. The method according to any one of claims 1 to 3, characterized in that: The step of inputting the mask sequence into the code generation model to obtain a first predicted code through the shared representation layer and the code understanding layer comprises: Determine a feature vector sequence corresponding to the mask sequence; For each feature vector in the feature vector sequence, the shared representation layer performs calculations based on all feature vectors before and after the feature vector in the feature vector sequence and the attention mechanism to obtain an intermediate feature vector; For each of the intermediate feature vectors, the code understanding layer performs attention calculation according to the intermediate feature vectors other than the intermediate feature vector and the attention mechanism to obtain a target feature vector; A first prediction program code is obtained according to the target feature vector.

5. The method according to any one of claims 1 to 3, characterized in that: The step of inputting the sample program code into the code generation model to obtain a second prediction code through the shared representation layer and the code generation layer comprises: Determine a feature vector sequence corresponding to the sample program code, wherein the feature vector sequence includes N feature vectors, where N is a positive integer; For the i-th feature vector in the feature vector sequence, the shared representation layer calculates according to all feature vectors before and after the i-th feature vector in the feature vector sequence and the attention mechanism to obtain an i-th intermediate feature vector, where i is a positive integer; For the i-th intermediate feature vector, the code generation layer performs calculation according to the preceding feature vector and the attention mechanism to obtain a target feature vector, wherein the preceding feature vector is all the intermediate feature vectors obtained before the i-th intermediate feature vector; A second prediction program code is obtained according to the target feature vector.

6. The method according to any one of claims 1 to 3, characterized in that: The sample program code is obtained in the following manner: Acquire a code data set, wherein the code data set includes a plurality of code files; For each of the code files, at least one of the following preprocessing steps is performed to obtain a training data set: when the ratio of the number of characters in the code file to the total number of symbols is greater than or equal to a preset character ratio, the code file is added to the training data set; when the average number of characters per line of code in the code file is less than or equal to a first preset threshold, the code file is added to the training data set; when the number of characters in the comment information in the code file is less than or equal to a second preset threshold, the code file is added to the training data set; The program code corresponding to each of the code files in the training data set is determined as a sample program code.

7. The method according to any one of claims 1 to 3, characterized in that: The code generation model is used to obtain the target program code generated based on the target text in the following manner: A predicted word is determined based on the target text, and the predicted word is used as the initial target word to loop through the following process: The target word is concatenated into the word segmentation sequence corresponding to the target text to obtain a target word segmentation sequence, a plurality of candidate prediction words and the probability of each candidate prediction word in a preset word list are determined based on the target word segmentation sequence, and a word is randomly selected from the candidate set corresponding to the candidate prediction word as a new target word until the length of the target word segmentation sequence reaches a preset length or the target word is a preset end symbol; Wherein, the sum of the probabilities of the candidate predicted words in the candidate set is greater than a preset probability.

8. A code generating device, characterized in that: The device comprises: An acquisition module, used for acquiring a target text, wherein the target text is a program code text to be supplemented or a natural language text for describing a code function; A generation module, used for inputting the target text into a code generation model to obtain a target program code generated based on the target text, wherein the code generation model includes a code understanding layer, a code generation layer and a shared representation layer, the result output by the shared representation layer is used to input the code understanding layer or the code generation layer, and the code generation model is trained by a code understanding task and a code generation task, the code understanding task is used for the code generation model to learn the grammatical features and semantic features of a sample program code, and the code generation task is used for the code generation model to learn the process of generating a new program code based on the sample program code; A training module is used to perform mask processing on the sample program code to obtain a mask sequence when executing the code understanding task, input the mask sequence into the code generation model to obtain a first prediction code through the shared representation layer and the code understanding layer, and adjust the parameters of the code understanding layer and the shared representation layer based on the first prediction code; when executing the code generation task, input the sample program code into the code generation model to obtain a second prediction code through the shared representation layer and the code generation layer, and adjust the parameters of the code generation layer and the shared representation layer based on the result output by the code generation model.

9. A non-transitory computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method described in any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Code generation method and system based on natural semantic understanding

    CN115202640A

  • Code completion model training method and code completion method and device

    CN115525263A