Code completion method
By converting code snippets into abstract syntax trees and inputting a tree-like BERT model, the problem that traditional BERT models cannot handle tree-like structure code is solved, and efficient code completion tasks are achieved.
Patent Information
- Application Number
- CN202510043205.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional BERT models perform well in natural language processing tasks, but cannot be applied directly to code completion tasks, because code snippets are usually tree-like structures rather than linear sequences.
Provide a code completion method. By converting the code snippet to be completed into an abstract syntax tree (AST), data preprocessing is performed on the AST, and the preprocessed data is obtained, and inputting it into the trained tree-like bidirectional encoding representation model based on the transformer, generating a code-completion vector sequence.
The code of the tree structure is converted into a processable form that the BERT model can be processed, so that the code completion can be effectively completed and the accuracy and efficiency of code completion are improved.
Smart Images

Figure CN119938011A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more specifically to a code completion method, device, equipment, medium and program product. Background Art
[0002] The main feature of the Bidirectional Encoder Representations from Transformers (BERT model) is that it uses the Transformer encoder to learn context-dependent word vectors. Transformer is a deep learning model with a self-attention mechanism that can capture long-range dependencies, allowing BERT to achieve remarkable results in natural language processing tasks.
[0003] The pre-training phase of the BERT model includes two tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP). In the MLM task, some words in the input text are randomly masked, and then the model needs to predict these masked words. In the NSP task, the model needs to determine whether two sentences are adjacent. Through the pre-training of these two tasks, the BERT model can learn rich semantic information and contextual representation.
[0004] Although the traditional BERT model performs well in natural language processing tasks, it cannot be directly applied to code completion tasks because code snippets are usually tree structures rather than linear sequences. Summary of the invention
[0005] In view of the above problems, the present disclosure provides a code completion method, apparatus, device, medium and program product.
[0006] According to a first aspect of the present disclosure, a code completion method is provided, including: in response to a code completion request, obtaining a code snippet to be completed; converting the code snippet to be completed into an abstract syntax tree; performing data preprocessing on the abstract syntax tree to obtain preprocessed data; and inputting the preprocessed data into a trained code completion model to generate a vector sequence of code completion, wherein the trained code completion model is trained by a tree-like transformer-based bidirectional encoding representation model.
[0007] According to an embodiment of the present disclosure, the method further includes: sorting or filtering the vector sequence for code completion to obtain the optimal vector sequence for code completion as the vector sequence for target code completion.
[0008] According to an embodiment of the present disclosure, the method further includes: converting the vector sequence of target code completion into a text form to obtain the completed code.
[0009] According to an embodiment of the present disclosure, converting a code snippet to be completed into an abstract syntax tree includes: using a code parser to convert the code snippet to be completed into an abstract syntax tree, wherein the abstract syntax tree includes a grammatical structure and an organizational relationship of the code snippet to be completed.
[0010] According to an embodiment of the present disclosure, data preprocessing is performed on an abstract syntax tree to obtain preprocessed data, including: mapping nodes of the abstract syntax tree to vector representations to obtain node vectors, wherein all node vectors constitute a vector sequence of a tree structure; traversing all node vectors, starting from the root node, and converting the vector sequence of the tree structure into a linear vector sequence; dividing the linear vector sequence into minimum units that can be processed by a tree-like transformer-based bidirectional coding representation model; mapping the minimum unit to a corresponding transformer-based bidirectional coding representation model vocabulary to obtain an input sequence; and adjusting the dimension of the input sequence so that the lengths of the input sequences are the same.
[0011] According to an embodiment of the present disclosure, a training method for a tree-like transformer-based bidirectional coding representation model includes: obtaining a pre-training sample data set, wherein the pre-training sample data set includes an input vector sequence suitable for the tree-like transformer-based bidirectional coding representation model; dividing the pre-training sample data set into a training set, a validation set, and a test set; based on a first loss function, training the tree-like transformer-based bidirectional coding representation model with the training set; and using an optimizer to perform back propagation and update the parameters of the tree-like transformer-based bidirectional coding representation model until the tree-like transformer-based bidirectional coding representation model converges to obtain a trained code completion model.
[0012] According to an embodiment of the present disclosure, the training method of the tree-shaped transformer-based bidirectional coding representation model also includes: verifying the trained code completion model with a verification set, selecting the model ranked first in performance as the optimal code completion model; testing the optimal code completion model with a test set and obtaining the test results; and evaluating the optimal code completion model based on the test results.
[0013] According to an embodiment of the present disclosure, the tree-shaped transformer-based bidirectional coding representation model is a pre-trained language model based on transformer-based bidirectional coding representation, and the tree-shaped transformer-based bidirectional coding representation model includes an input layer and a multi-layer Transformer encoder.
[0014] The second aspect of the present disclosure provides a code completion device, including: an acquisition module, used to respond to a code completion request and obtain a code snippet to be completed; a conversion module, used to convert the code snippet to be completed into an abstract syntax tree; a preprocessing module, used to perform data preprocessing on the abstract syntax tree to obtain preprocessed data; and a generation module, used to input the preprocessed data into a trained code completion model to generate a vector sequence for code completion, wherein the trained code completion model is trained by a tree-like transformer-based bidirectional encoding representation model.
[0015] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0016] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the above computer program or instructions are executed by a processor.
[0017] The fifth aspect of the present disclosure further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when the above computer program or instructions are executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0019] Figure 1 The application scenario diagram of the code completion method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown;
[0020] Figure 2 A flowchart of a method for training a tree-like BERT model according to an embodiment of the present disclosure is schematically shown;
[0021] Figure 3 A flowchart of a code completion method according to an embodiment of the present disclosure is schematically shown;
[0022] Figure 4 A flowchart schematically shows a method of performing data preprocessing on an abstract syntax tree to obtain preprocessed data according to an embodiment of the present disclosure;
[0023] Figure 5 A structural block diagram of a code completion device according to an embodiment of the present disclosure is schematically shown; and
[0024] Figure 6A block diagram of an electronic device suitable for implementing a code completion method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0027] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0028] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0029] In the related art, when performing computer network programming or project programming, programmers often use intelligent code completion methods to assist programming. Existing intelligent code completion includes the following types: The first type is based on the latest open source deep learning model, which trains a model suitable for code text completion based on historical massive open source code. In the prediction stage of actual application, the model reads the context to predict the list of completion candidates for the current cursor position and displays it to the user. The deep learning model can complete longer code statements, but the accuracy is low, and it often recommends code that is not suitable for the context or has grammatical errors, which causes trouble to users. The second type is to make predictions based on grammatical information. This code completion method does not have a deep learning model prediction process. It generally recommends the smallest unit (token) into which a single text data is divided (), which is fast but short in length. It is a built-in capability of the code editor and cannot achieve intelligent code completion. Current code completion methods basically rely on deep learning models, but the models are relatively black-box. Models based on historical training data are prone to recommending code that does not conform to the contextual grammar, forcing users to go back and modify large sections of completed code, which causes trouble to users and greatly reduces the efficiency of completion tools.
[0030] An embodiment of the present disclosure provides a code completion method, including: in response to a code completion request, obtaining a code snippet to be completed; converting the code snippet to be completed into an abstract syntax tree; performing data preprocessing on the abstract syntax tree to obtain preprocessed data; and inputting the preprocessed data into a trained code completion model to generate a vector sequence of code completion, wherein the trained code completion model is trained by a tree-like bidirectional encoder representation model (Bidirectional Encoder Representations from Transformers, BERT model).
[0031] On the other hand, an embodiment of the present disclosure provides a training method for a tree-like BERT model, including: obtaining a pre-training sample data set, wherein the pre-training sample data set includes an input vector sequence suitable for the tree-like BERT model; dividing the pre-training sample data set into a training set, a validation set, and a test set; based on a first loss function, training the tree-like BERT model with the training set; and using an optimizer to perform backpropagation and update the parameters of the tree-like BERT model until the tree-like BERT model converges to obtain a trained code completion model.
[0032] Figure 1 The application scenario diagram of the code completion method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.
[0033] like Figure 1As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0034] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0035] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0036] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0037] It should be noted that the code completion method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the code completion device provided in the embodiment of the present disclosure can generally be set in the server 105. The code completion method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Correspondingly, the code completion device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.
[0039] The following will be based on Figure 1 The scene described by Figure 2~Figure 4 The code completion method of the disclosed embodiment is described in detail.
[0040] Although traditional BERT performs well in natural language processing tasks, it cannot be directly applied to code completion tasks because code snippets are usually tree structures rather than linear sequences. In order to process tree-structured code, the present disclosure uses a tree-like BERT to construct an abstract syntax tree (AST). The tree-like BERT can represent code snippets in the form of AST, which is the result of the editor's first step of processing the code. It is a source code represented in the form of a tree, and each element of the source code is mapped to a node or subtree.
[0041] The Tree-BERT model is an extension of the traditional BERT model, specifically designed to process tree-structured data, such as code snippets. The model structure of the Tree-BERT is similar to that of BERT, but it converts the input sequence into a tree-like hierarchical interface for encoding. The input sequence is converted into a tree, and then the nodes of the tree are converted into sequences and input into the BERT model for encoding.
[0042] Tree-like BERT is a deep learning model, so the training process may require a lot of computing resources and time, and is usually accelerated using graphics processors or tensor processors.
[0043] Figure 2 The flowchart of the training method of the tree-like BERT model according to an embodiment of the present disclosure is schematically shown.
[0044] like Figure 2 As shown, the training method of the tree-like BERT model of this embodiment includes operations S210 to S240.
[0045] In operation 210, a pre-training sample dataset is obtained, wherein the pre-training sample dataset includes an input vector sequence suitable for a tree-structured BERT model.
[0046] In the embodiment of the present disclosure, the tree-like BERT model is a pre-trained language model based on BERT. The tree-like BERT model includes an input layer and a multi-layer Transformer encoder (the Transformer encoder is responsible for encoding the original text data into an intermediate state vector). The parameter settings of the tree-like BERT are similar to those of BERT, including: the following window size is 512. Number of layers: 12 layers, each with 768 hidden units. The embedding vector size is 768. The learning rate is 2e -5 .
[0047] In an embodiment of the present disclosure, the pre-training sample data set can be obtained by preprocessing a test case data set suitable for code completion. The data set should contain some incomplete code snippets and corresponding expected outputs, that is, complete code snippets, including test cases, incomplete test case codes, and code snippets that can be completed as completed test cases. The quality and scale of the data set are crucial to the training and performance of the model. When constructing the data set, it is necessary to determine the length of the code snippet. Specifically, according to the task requirements and the input restrictions of the model, the length of the code snippet can be limited to facilitate the input of the tree-like BERT model. The test case data set preprocessing process includes: converting the code snippet into an abstract syntax tree (AST). AST is generated by a code parser, which captures the grammatical structure and organizational relationship of the code; mapping each node of AST to a vector representation, and methods such as word embedding can be used; starting from the root node, constructing an input vector sequence, and expanding the tree structure into a linear sequence. Methods such as pre-order traversal or post-order traversal can be used; dividing the input sequence into the smallest unit (token) that the model can process, and mapping it to the corresponding BERT vocabulary; and padding the sequence so that all input sequences have the same length for batch processing. Among them, word embedding is a technology in natural language processing (NLP) that maps words into real vector space so that semantically similar words are close to each other in the vector space. This method can capture rich relationships between words, including synonyms, antonyms, hyponyms, and hyponyms.
[0048] In operation 220 , the pre-training sample dataset is divided into a training set, a validation set, and a test set.
[0049] In an embodiment of the present disclosure, a data set is divided into a training set, a validation set, and a test set for model training, tuning, and evaluation. The training set, validation set, and test set can be divided according to different preset ratios to achieve different purposes of iterative training. For example, the training set, validation set, and test set are divided according to 7:2:1.
[0050] In operation 230 , a tree-structured BERT model is trained using the training set based on the first loss function.
[0051] In an embodiment of the present disclosure, the first loss function is a loss function suitable for code completion tasks, such as mean square error loss or cross entropy loss. The loss function is a function used to measure the degree of inconsistency between the predicted value of the model and the true value. In machine learning, it is hoped that the predicted value is infinitely close to the true value, so the difference needs to be minimized, which requires the introduction of a loss function. The smaller the loss function, the better the robustness of the model. Among them, mean square error loss (MSE) is a loss function commonly used in regression tasks, which calculates the average of the sum of squares of the differences between the model predicted value and the actual target value; cross entropy describes the distance between two probability distributions, that is, the smaller the cross entropy value (the smaller the relative entropy value), the closer the two probability distributions are.
[0052] In operation 240 , the optimizer is used to perform back propagation and update the parameters of the tree-like BERT model until the tree-like BERT model converges to obtain a trained code completion model.
[0053] In the embodiments of the present disclosure, a suitable optimizer is selected to update the model parameters to minimize the loss function. The first-order moment estimate (mean of the gradient) and the second-order moment estimate (uncentered variance of the gradient) of the gradient are comprehensively considered to calculate the update step size. It has the following advantages: simple implementation, efficient calculation, and low memory requirement; parameter update is not affected by the scaling transformation of the gradient; hyperparameters have good interpretability and usually do not need to be adjusted or only require very little fine-tuning; the update step size can be limited to an approximate range (initial learning rate); the step size annealing process can be naturally implemented (automatically adjusting the learning rate); suitable for applications in large-scale data and parameter scenarios; suitable for unstable objective functions; and suitable for problems where the gradient is sparse or there is a lot of noise in the gradient.
[0054] In the embodiment of the present disclosure, the learning process of the back propagation algorithm consists of a forward propagation process and a back propagation process. In the forward propagation process, the input information passes through the input layer and the hidden layer, is processed layer by layer and transmitted to the output layer. If the expected output value is not obtained in the output layer, the sum of the squares of the error between the output and the expected value is taken as the objective function, and the back propagation is turned to, and the partial derivatives of the objective function to the weights of each neuron are obtained layer by layer to form the gradient of the objective function to the weight vector. As the basis for modifying the weights, the learning of the network is completed in the weight modification process. When the error reaches the expected value, the network learning ends. The back propagation algorithm can avoid repeated calculations and improve training efficiency.
[0055] In an embodiment of the present disclosure, in order to ensure the performance and effect of the tree-like BERT model in the method, the model can be evaluated and tuned. Some metrics, such as accuracy, recall, F1 value (F1 value is an indicator to measure the accuracy of a binary or multi-task binary classification model. It is the harmonic mean of precision and recall.) can be used to evaluate the performance of the model. If the model effect is not ideal, you can try to adjust the model's hyperparameters, the construction of the data set, increase the size of the data set, modify the AST representation method, and other means to improve the model's code completion performance. The model training method may also include: verifying the trained code completion model with a validation set, selecting the model ranked first in performance as the optimal code completion model; testing the optimal code completion model with a test set, and obtaining the test results; and evaluating the optimal code completion model based on the test results.
[0056] For example, the n tree-like BERT models trained with the validation set are verified, the loss value is calculated using the first loss function, and the trained tree-like BERT model with the smallest loss value is selected as the optimal tree-like BERT model. The model is then tested with the test set data, and the performance of the model is evaluated based on the loss value of the test results. The smaller the loss value, the better the performance of the model. If the performance of the model after evaluation does not meet the expected requirements, the model parameters can be adjusted, and the training set data can be added to further train and test the model.
[0057] Figure 3 The flowchart of the code completion method according to the embodiment of the present disclosure is schematically shown.
[0058] like Figure 3 As shown, the code completion method of this embodiment includes operations S310 to S340.
[0059] In operation 310 , in response to a code completion request, a code segment to be completed is obtained.
[0060] In an embodiment of the present disclosure, the code snippet to be completed includes an incomplete test case to be completed. According to task requirements and model input restrictions, the length of the code snippet can be limited to facilitate input into the tree-like BERT model.
[0061] In operation 320 , the code snippet to be completed is converted into an abstract syntax tree.
[0062] In an embodiment of the present disclosure, for a code completion task, it is necessary to represent the code snippet as a tree structure so that the tree-like BERT model can be applied. The code snippet to be completed is converted into an abstract syntax tree using a code parser, wherein the abstract syntax tree includes the grammatical structure and organizational relationship of the code snippet to be completed. AST is generated by a code parser to capture the code grammatical structure and code organizational relationship of the code. Representing the code as AST can better capture the structural information of the code, thereby improving the accuracy and effect of the code completion task.
[0063] There may be different AST representation methods for different programming languages, so you need to select and build AST according to the specific programming language. Some programming languages provide ready-made AST libraries that can be used directly. For some special languages or situations where custom AST representation is required, you need to write your own code parser to generate AST.
[0064] In operation 330, data preprocessing is performed on the abstract syntax tree to obtain preprocessed data.
[0065] In an embodiment of the present disclosure, the tree-like BERT model needs to represent the input as a vector sequence, and the input needs to be decomposed and filled. In the code completion task, a pre-trained tree-like BERT model can be used, and fine-tuned on its basis to adapt to the specific code completion task. Among them, the purpose of the decomposition operation is to divide the input text into tokens, and cooperate with the dictionary to allow the machine to recognize the text; filling is to add additional boundary values around the input data. These boundary values are usually some specific values, such as 0. In convolutional neural networks, the role of padding is to adjust the dimension of the input data to ensure that the size of the feature map after the convolution or pooling operation is the same as the input feature map. This can keep the number of network parameters and the amount of computation within an acceptable range, and also help to improve the generalization ability of the network.
[0066] Figure 4 The flowchart schematically shows a method of performing data preprocessing on an abstract syntax tree to obtain preprocessed data according to an embodiment of the present disclosure.
[0067] like Figure 4 As shown, the method of performing data preprocessing on the abstract syntax tree to obtain preprocessed data in this embodiment includes operations S331 to 335.
[0068] In operation S331, the nodes of the abstract syntax tree are mapped to vector representations to obtain node vectors, wherein all node vectors constitute a vector sequence of a tree structure.
[0069] In the embodiments of the present disclosure, in order to make the obtained node vectors relatively complete, multiple convenient methods can be used. For example, an abstract syntax tree to be processed is obtained; a breadth-first traversal is performed on the abstract syntax tree to obtain a first sequence, and a depth-first traversal is performed on the abstract syntax tree to obtain a second sequence; a coding sequence to be processed is generated according to the first sequence and the second sequence; the coding sequence to be processed is processed by a pre-constructed vectorized processing model to obtain a vectorized representation result of the nodes in the abstract syntax tree.
[0070] In operation S332, all node vectors are traversed, starting from the root node, and the vector sequence of the tree structure is converted into a linear vector sequence.
[0071] In an embodiment of the present disclosure, a directed edge set corresponding to a data source to be processed is obtained, and each directed edge in the directed edge set is traversed to determine a first root node; each directed edge in the directed edge set is traversed according to the first root node to determine a child node of the first root node; the child node is determined as a second root node, and each directed edge in the directed edge set is repeatedly traversed according to the second root node to determine a child node of the second root node, until the traversal result of the directed edge meets a preset traversal condition, and data in a tree structure is constructed based on the determined root node and child nodes.
[0072] In operation S333, the linear vector sequence is divided into tokens that can be processed by the tree-like BERT model.
[0073] In the embodiments of the present disclosure, a token refers to the basic data unit processed by the model. It can be a word, a character, a phrase, or even an image fragment, a sound fragment, etc. For example, a sentence will be divided into multiple tokens, and each punctuation mark will also be regarded as a separate token. The way tokens are divided will affect the model's understanding and processing of data.
[0074] In operation S334, the token is mapped to the corresponding BERT vocabulary to obtain an input sequence.
[0075] In operation S335 , the dimensions of the input sequences are adjusted so that the lengths of the input sequences are the same.
[0076] In an embodiment of the present disclosure, each node of the AST is mapped to a vector representation, and methods such as word embedding can be used. Then, the input vector sequence is constructed in a pre-order traversal manner: ["@Test", "testAddition", "Block", "int", "result", "=", "add", "result", "=", "add", "result", "2", "3"]. Next, the input sequence is decomposed and padded to ensure that its length is the same, so as to facilitate the input of the tree-like BERT model.
[0077] Return to reference Figure 3 In operation 340, the preprocessed data is input into the trained code completion model to generate a vector sequence of code completion, wherein the trained code completion model is trained by the tree-structured BERT model.
[0078] In an embodiment of the present disclosure, the code completion method may further include: sorting or filtering the vector sequence of code completion to obtain the optimal vector sequence of code completion as the vector sequence of target code completion. For example, some sorting or filtering algorithms may be used to sort or filter the prediction results to improve their accuracy and usability.
[0079] In an embodiment of the present disclosure, the code completion method may further include: converting a vector sequence of target code completion into a text form to obtain a completed code.
[0080] In the embodiments of the present disclosure, a tree-like BERT model structure is established through a pre-trained model BERT to establish an abstract syntax tree of the code, and the constructed tree-like BERT model is trained by using the asset code in the code repository. After the model training is completed, the trained tree-like BERT model is used to complete the specific code. By completing the code by itself through the tree-like BERT, the coding workload of the developer can be reduced. On the other hand, the code can be detected, and the missing statement information in the test case can be output, and then the code completion evaluation can be performed.
[0081] For example, in a complex application project, developers need to write a lot of backend logic code to handle user requests and data interactions. Through the tree-like BERT model, you only need to enter part of the code framework or key function name, and the tree-like BERT model can predict and complete the remaining code structure, including variable declarations, conditional judgments, loop structures, etc.
[0082] In addition, the tree-like BERT model can also perform in-depth analysis of existing code to identify potential logic loopholes or uncovered test scenarios. For example, in the transaction processing module of a certain platform, developers may have omitted the transaction status update logic under certain specific conditions. By traversing the code structure, the tree-like BERT model can accurately locate these missing statements and output a detailed report indicating which test cases failed to cover these logical branches. Developers can quickly locate problems based on the report and perform targeted code completion and optimization, thereby improving the overall quality and stability of the code.
[0083] In addition, the tree-like BERT model can also assist in evaluating the effectiveness of code completion. After completing the code, the tree-like BERT model can re-analyze the code, compare the code structure and logical integrity before and after completion, and give corresponding scores or suggestions. This intelligent evaluation method not only improves the efficiency of code review, but also ensures the accuracy and effectiveness of code completion.
[0084] Based on the above code completion method, the present disclosure also provides a code completion device. Figure 5 The device is described in detail.
[0085] Figure 5 The structural block diagram of the code completion device according to the embodiment of the present disclosure is schematically shown.
[0086] like Figure 5 As shown, the code completion device 800 of this embodiment includes an acquisition module 810, a conversion module 820, a preprocessing module 830 and a generation module 840.
[0087] The acquisition module 810 is used to respond to the code completion request and acquire the code snippet to be completed. In one embodiment, the acquisition module 810 can be used to perform the operation S310 described above, which will not be described in detail here.
[0088] The conversion module 820 is used to convert the code snippet to be completed into an abstract syntax tree. In one embodiment, the conversion module 820 can be used to perform the operation S320 described above, which will not be described in detail here.
[0089] The preprocessing module 830 is used to perform data preprocessing on the abstract syntax tree to obtain preprocessed data. In one embodiment, the preprocessing module 830 can be used to perform the operation S330 described above, which will not be described in detail here.
[0090] The generation module 840 is used to input the preprocessed data into the trained code completion model to generate a vector sequence for code completion, wherein the trained code completion model is obtained by training the tree-like BERT model. In one embodiment, the generation module 840 can be used to perform the operation S340 described above, which will not be repeated here.
[0091] According to an embodiment of the present disclosure, the code completion apparatus 800 includes an acquisition module 810 , a conversion module 820 , a preprocessing module 830 and a generation module 840 .
[0092] According to an embodiment of the present disclosure, the preprocessing module 830 may include a first conversion submodule 831 , a second conversion submodule 832 , a third conversion submodule 833 , a fourth conversion submodule 834 , and a fifth conversion submodule 835 .
[0093] The first conversion submodule 831 is used to map the nodes of the abstract syntax tree to vector representations to obtain node vectors, wherein all node vectors constitute a vector sequence of a tree structure. In one embodiment, the first conversion submodule 831 can be used to perform the operation S330 described above, which will not be repeated here.
[0094] The second conversion submodule 832 traverses all node vectors, starting from the root node, and converts the vector sequence of the tree structure into a linear vector sequence. In one embodiment, the second conversion submodule 832 can be used to perform the operation S332 described above, which will not be described in detail here.
[0095] The third conversion submodule 833 is used to divide the linear vector sequence into tokens that can be processed by the tree-like BERT model. In one embodiment, the third conversion submodule 833 can be used to perform the operation S333 described above, which will not be repeated here.
[0096] The fourth conversion submodule 834 is used to map the token to the corresponding BERT vocabulary to obtain an input sequence. In one embodiment, the fourth conversion submodule 834 can be used to perform the operation S334 described above, which will not be described in detail here.
[0097] The fifth conversion submodule 835 is used to adjust the dimension of the input sequence so that the lengths of the input sequences are the same. In one embodiment, the fifth conversion submodule 835 can be used to perform the operation S335 described above, which will not be described in detail here.
[0098] According to an embodiment of the present disclosure, any multiple modules of the acquisition module 810, the conversion module 820, the preprocessing module 830 and the generation module 840 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the acquisition module 810, the conversion module 820, the preprocessing module 830 and the generation module 840 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in a suitable combination of any of them. Alternatively, at least one of the acquisition module 810 , the conversion module 820 , the preprocessing module 830 , and the generation module 840 may be at least partially implemented as a computer program module, and when the computer program module is executed, a corresponding function may be performed.
[0099] Figure 6 A block diagram of an electronic device suitable for implementing a code completion method according to an embodiment of the present disclosure is schematically shown.
[0100] like Figure 6 As shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 to a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include an onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0101] In RAM 903, various programs and data required for the operation of electronic device 900 are stored. Processor 901, ROM 902 and RAM 903 are connected to each other via bus 904. Processor 901 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 902 and / or RAM 903. It should be noted that the program can also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in one or more memories.
[0102] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the input / output (I / O) interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage portion 908 as needed.
[0103] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0104] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.
[0105] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the code completion method provided by the embodiment of the present disclosure.
[0106] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 901. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0107] In one embodiment, the computer program may be based on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0108] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0109] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0110] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0111] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0112] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A code completion method, characterized in that: The method comprises: In response to the code completion request, obtaining a code snippet to be completed; Convert the code snippet to be completed into an abstract syntax tree; Performing data preprocessing on the abstract syntax tree to obtain preprocessed data; and Input the preprocessed data into the trained code completion model to generate a vector sequence for code completion, The trained code completion model is obtained by training a tree-like transformer-based bidirectional coding representation model.
2. The method according to claim 1, characterized in that The method further includes: sorting or filtering the code completion vector sequence to obtain an optimal code completion vector sequence as a target code completion vector sequence.
3. The method according to claim 2, characterized in that The method further comprises: The vector sequence of the target code completion is converted into a text form to obtain the completed code.
4. The method according to claim 1, characterized in that: The converting the code snippet to be completed into an abstract syntax tree includes: using a code parser to convert the code snippet to be completed into an abstract syntax tree, wherein the abstract syntax tree includes the grammatical structure and organizational relationship of the code snippet to be completed.
5. The method according to claim 1, characterized in that Performing data preprocessing on the abstract syntax tree to obtain preprocessed data includes: Mapping the nodes of the abstract syntax tree to vector representations to obtain node vectors, wherein all the node vectors constitute a vector sequence of a tree structure; Traversing all the node vectors, starting from the root node, converting the vector sequence of the tree structure into a linear vector sequence; Dividing the linear vector sequence into minimum units that can be processed by the tree-like transformer-based bidirectional coding representation model; Mapping the minimum unit into a corresponding transformer-based bidirectional encoding representation model vocabulary to obtain an input sequence; and The dimensions of the input sequences are adjusted so that the lengths of the input sequences are the same.
6. The method according to claim 1, characterized in that The training method of the tree-shaped transformer-based bidirectional encoding representation model includes: Acquire a pre-training sample data set, wherein the pre-training sample data set includes an input vector sequence suitable for a tree-like transformer-based bidirectional coding representation model; Dividing the pre-training sample data set into a training set, a validation set and a test set; Based on a first loss function, training the tree-like transformer-based bidirectional encoding representation model with the training set; and The optimizer is used to perform back propagation and update the parameters of the tree-like transformer-based bidirectional coding representation model until the tree-like transformer-based bidirectional coding representation model converges to obtain a trained code completion model.
7. The method according to claim 6, characterized in that The training method of the tree-shaped transformer-based bidirectional coding representation model also includes: Using the validation set to validate the trained code completion model, and selecting the model ranked first in performance as the optimal code completion model; Using the test set to test the optimal code completion model, and obtaining the test result; The optimal code completion model is evaluated according to the test results.
8. The method according to any one of claims 1 to 7, characterized in that The tree-shaped transformer-based bidirectional coding representation model is a pre-trained language model based on transformer-based bidirectional coding representation, and the tree-shaped transformer-based bidirectional coding representation model includes an input layer and a multi-layer Transformer encoder.
9. A code completion device, characterized in that: The device comprises: An acquisition module, used to respond to a code completion request and acquire a code snippet to be completed; A conversion module, used for converting the code snippet to be completed into an abstract syntax tree; A preprocessing module, used to perform data preprocessing on the abstract syntax tree to obtain preprocessed data; and A generation module is used to input the preprocessed data into the trained code completion model to generate a vector sequence for code completion, The trained code completion model is obtained by training a tree-like transformer-based bidirectional coding representation model.
10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.