Domain specific code completion method based on model collaborative reasoning
By fine-tuning the smaller-scale language models and building task-specific classifiers, the problem of large language models performing poorly in domain-specific code completion is solved, and high-precision code completion and cost-reducing effects are achieved.
Patent Information
- Application Number
- CN202510115286.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing large language models perform poorly when completing domain-specific code completion tasks because they do not necessarily learn domain-specific code knowledge and are difficult to capture the dependencies between domain-specific code structures and code.
By fine-tuning the smaller-scale language model, learning domain-specific code knowledge, and building task-specific classifiers, extracting multi-dimensional features related to code completion to assist the classifier in automatically fusing the inference results of different language models.
High-precision code completion is achieved, reducing the cost of fine-tuning the model, and improving the inference accuracy.
Smart Images

Figure CN120066569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of code completion, and in particular to a domain-specific code completion method based on model collaborative reasoning. Background Art
[0002] Code completion is an important task in the field of software engineering, aiming to use code completion tools to complete code based on the existing code context, which can help developers write code quickly and automatically, improving development efficiency. Currently, researchers have proposed many methods for automatically completing code to improve development efficiency and ensure software quality.
[0003] Most of the current code completion methods are built based on large language models. Large language models such as Deepseek-Coder, CodeLlama, and StarCoder are trained on a large amount of code corpora and have strong code reasoning capabilities, which can help developers efficiently complete high-quality code. However, these large language models perform poorly when completing domain-specific code completion tasks because these models are trained on general code and may not learn domain-specific code knowledge, making it difficult to capture the dependencies between domain-specific code structures and code.
[0004] Although existing research work has started to improve the code reasoning ability of large language models in specific domains through fine-tuning techniques or retrieval enhancement techniques, etc., the cost of fine-tuning large language models is extremely high, and the retrieval enhancement technique has encountered challenges in how to retrieve effective information and how to correctly guide the large language model to use the retrieved information to correctly complete code completion, etc. Different from them, the present invention aims to learn domain-specific code knowledge only by fine-tuning a relatively small-scale language model, thereby reducing the cost of the fine-tuned model, and at the same time constructing a task-specific classifier to extract multi-dimensional features related to code completion to assist the classifier in automatically fusing the reasoning results of different language models, so as to achieve high-precision code completion. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a domain-specific code completion method based on model collaborative reasoning.
[0006] The purpose of the present invention is achieved through the following technical solutions: A domain-specific code completion method based on model collaborative reasoning, the method includes the following steps:
[0007] Collect the code of open-source projects for a specific domain, construct a domain-specific code completion dataset, and divide the dataset into a training set, a validation set, and a test set; use the code in the training set to fine-tune a small-scale language model using the PEFT method, so that the language model learns domain-specific code knowledge;
[0008] Extract the code from the training set and the validation set, construct line-level code completion examples respectively, and use the large language model and the fine-tuned small-scale language model to reason about the constructed examples respectively to complete the code completion task; collect the reasoning results of the large language model and the fine-tuned small-scale language model during the reasoning process, and retrieve from the context of the input code to be completed and the code training set according to the outputs of the large language model and the fine-tuned small-scale language model to extract multi-dimensional features, and construct the training set and the validation set for training the domain-specific classifier;
[0009] Build a model collaborative reasoning framework, and complete the reasoning in combination with the speculative decoding algorithm. The specific reasoning process is as follows: first, the fine-tuned small-scale language model performs serial reasoning, and then the large language model performs parallel generation based on the reasoning results of the small-scale language model to improve the reasoning speed of the large language model; after both the large language model and the fine-tuned small-scale language model complete the reasoning, select the parts with differences in the output codes of the large language model and the fine-tuned small-scale language model using the classifier model to fuse the reasoning results of the large language model and the fine-tuned small-scale language model and improve the reasoning accuracy.
[0010] Further, the number of parameters of the small-scale language model is within 2B.
[0011] Further, the multi-dimensional features include the model output dimension, the code type dimension, the code length dimension, the code frequency dimension, the code similarity dimension, and the retrieval information dimension;
[0012] The model output dimension includes the following features:
[0013] The input vector of the last layer of the model: During the model prediction process, the vector input to the last layer of each language model;
[0014] The probability output by the model: During the model prediction process, the probability that each language model outputs the specified code;
[0015] The code type dimension includes the following features:
[0016] Code type: The code type output by each language model, including line breaks, punctuation, keywords, identifiers, numbers, and the EOS special token output by the model;
[0017] The code length dimension includes the following features:
[0018] The length of the currently completed code: The length of the code that the current model has completed;
[0019] The length of the input code: The length of the code input to the model (i.e., the context of the line to be completed);
[0020] The code frequency dimension includes the following features:
[0021] The occurrence frequency of the current completed code: the occurrence frequency of all the codes completed by the current model in the context of the code input to the model;
[0022] The occurrence frequency of the current completed line: the occurrence frequency of the lines composed of all the codes completed by the current model in the context of the code input to the model;
[0023] The code similarity dimension includes the following features:
[0024] The similarity of the current completed line: the similarity between the line composed of all the codes completed by the current model and the context of the code input to the model;
[0025] The similarity of the current completed code snippet: the similarity between the line composed of all the codes completed by the current model and the code snippet formed by the context of the line to be completed input to the model, calculating the similarity between this code snippet and the code context;
[0026] The retrieval information dimension includes the following features:
[0027] The occurrence frequency of the current completed line in the training set: the occurrence frequency of the line composed of all the codes completed by the current model in the code training set;
[0028] The occurrence frequency of the current completed code snippet in the training set, the line composed of all the codes completed by the current model and the context of the completed line form a code snippet, calculating the occurrence frequency of this code snippet in the training set;
[0029] The similarity between the current completed line and the training set: the similarity between the line composed of all the codes completed by the current model and the code training set;
[0030] The similarity between the current completed code snippet and the training set: the line composed of all the codes completed by the current model and the context of the completed line form a code snippet, calculating the similarity between this code snippet and the training set.
[0031] Furthermore, a model collaborative inference framework is constructed, and the specific inference is completed by combining the speculative decoding algorithm as follows:
[0032] Use the fine-tuned small-scale language model for serial inference to complete code completion; the input is the context c of the code to be completed t =(x 1 ,x 2 ,…,x t-1 ), and the output is the candidate code sequence S=(y t ,y t+1 ,…,y t+l-1 ), where x 1, x 2 , …, x t-1 and y t , y t+1 , …, y t+l-1 represent the tokens corresponding to the code, and l represents the length of the code output by the fine-tuned small-scale language model;
[0033] According to the input code context and the candidate code generated by the fine-tuned model, use the large language model for parallel inference to complete code completion; the input is (x 1 , x 2 , …, x t-1 , y t , y t+1 , …, y t+l-1 ), and the output is the candidate code sequence S ′ = (y t ′ , y t ′ +1 , …, y t ′ +l-1 ), where y t ′ etc. represent the tokens corresponding to the code output by the large language model;
[0034] Use a classifier to fuse S and S ′: Starting from position t, judge one by one. If the codes output by the large language model and the fine-tuned small-scale language model at the same position j are different, y j ≠ y j ′ , extract features according to the output information when the fine-tuned small-scale language model generates the corresponding code and input them to the classifier, and the classifier makes a selection to obtain the complete output code sequence (y t , y j-1 , y j ′ ).
[0035] The beneficial effects of the present invention are as follows: The present invention fine-tunes a small-scale language model to learn domain-specific code knowledge, effectively reducing the model deployment and training costs compared with directly fine-tuning a large-scale language model; extracting features from the outputs of the large language model and the fine-tuned language model to train a classifier can help adaptively and efficiently combine the inference results of the large language model and the fine-tuned language model, improving the inference accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the overall framework diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The present invention will be further described below with reference to the accompanying drawings and examples.
[0038] As shown Figure 1 in the figure, the embodiment of the present invention provides a domain-specific code completion method based on model collaborative reasoning, which specifically includes the following steps:
[0039] Step 1. Collect the code of open-source projects for a specific domain, construct a dataset, and divide it into a training set, a validation set, and a test set. According to the code in the training set, use the PEFT method to fine-tune a relatively small language model, reducing the fine-tuning cost while enabling the model to learn domain-specific code knowledge.
[0040] Step 2. Extract the code from the training set, construct line-level code completion test examples, and use a large language model and the fine-tuned language model to perform reasoning on the constructed examples respectively to complete the code completion task. During the process, collect the reasoning results of different language models, extract multi-dimensional features based on the outputs of each language model, and use the multi-dimensional features to construct a classifier model using a two-layer MLP architecture. The multi-dimensional features include:
[0041] a. Model output dimension. The features of this dimension are extracted based on the output of the model itself and include the following two specific features:
[0042] The input vector EMB of the last layer of the model: During the model prediction process, the vector input to the last layer lm_head of the language model contains the information refined through the layers of the model.
[0043] The probability PT of the model output: During the model prediction process, the probability that the language model outputs the specified code. The higher the probability, the higher the confidence of the model in the currently output code, and the more credible the currently output code is to a certain extent.
[0044] b. Code type dimension TW. The features of this dimension are extracted based on the code output by the model at the current position. In the present invention, the code types are divided into line breaks, punctuation marks, keywords, identifiers, numbers, and special tokens such as EOS output by the model.
[0045] c. Code length dimension. The features of this dimension are extracted based on all the code already output by the current model and the code context input to the model, and are used to assist other features in making judgments. It includes the following two specific features:
[0046] The length LW of the currently completed code: The length of the code already completed by the current model.
[0047] The length LL of the context above the currently completed line: The number of lines of the code context input to the current model.
[0048] d. Code frequency dimension. The features in this dimension are calculated based on all the codes that have been output by the current model and the code context of the input model, and include the following two specific features:
[0049] The occurrence frequency FW of the current completed code: The occurrence frequency of the latest code completed by the current model in the code context. Considering the possible repetition of the code itself, such as previously defined variables may be used in subsequent codes, it is considered that the higher the code occurrence frequency, the higher the credibility.
[0050] The occurrence frequency FL of the current completed line: The occurrence frequency of the line composed of all the codes that have been completed by the current model in the code context. The higher the occurrence frequency, the higher the credibility.
[0051] e. Code similarity dimension. The features in this dimension are calculated based on all the codes that have been output by the current model and the code context of the input model, and include the following two specific features:
[0052] The similarity SL of the current completed line: The BM25 similarity between the line composed of all the codes that have been completed by the current model and the code context. The higher the similarity, the greater the correlation between the current completed content and the code context, and the higher the credibility.
[0053] The similarity SS of the current completed code snippet: The line composed of all the codes that have been completed by the current model and the code snippet composed of the context above the completed line. Calculate the BM25 similarity between this code snippet and the code context. The higher the similarity, the greater the correlation between the current completed content and the code context, and the sufficient correlation with the context above the line to be completed, and the higher the credibility.
[0054] f. Retrieval information dimension. The features in this dimension are calculated by retrieving and calculating all the codes that have been output by the current model and the code context of the input model in the code library, and include the following four specific features:
[0055] The occurrence frequency RFL of the current completed line in the training set: The occurrence frequency of the line composed of all the codes that have been completed by the current model in the code training set. The higher the occurrence frequency, the higher the credibility.
[0056] The occurrence frequency RFS of the current completed code snippet in the training set: The line composed of all the codes that have been completed by the current model and the code snippet composed of the context above the completed line. Calculate the occurrence frequency of this code snippet in the training set. The higher the occurrence frequency, the higher the credibility.
[0057] The similarity RSL of the current completed line and the training set: The Jaccard similarity between the line composed of all the codes that have been completed by the current model and the code training set. The higher the similarity, the greater the correlation between the current completed content and the code library, that is, the stronger the correlation with the current field, and the higher the credibility.
[0058] The similarity RSS between the current code completion snippet and the training set: The lines composed of all the code completed by the current model and the code snippet formed by the context above the completed line are used to calculate the Jaccard similarity between this code snippet and the training set. The higher the similarity, the greater the correlation between the current completion content and the code library, that is, the stronger the correlation with the current field and the higher the credibility.
[0059] After extracting all features, input them to the two-layer MLP architecture to train the classifier model.
[0060] Step 3. Build a model collaborative inference framework, complete the inference by combining the speculative decoding algorithm, and use the classifier to fuse the inference results of different language models. The specific inference steps are as follows:
[0061] 3.1. Use a smaller language model that has been fine-tuned for serial inference to complete code completion. Input the context c of the code to be completed t =(x 1 ,x 2 ,…,x t-1 ), call the fine-tuned language model to generate code y t , and then splice to obtain the new code context c t+1 =(x 1 ,x 2 ,…,x t-1 ,y t ), continue to input it to the model to generate code y t+1 , and so on in a loop until l codes are generated to obtain the candidate code sequence S=(y t ,y t+1 ,…,y t+l-1 ).
[0062] 3.2. Use a large language model for parallel inference based on the input code context and the candidate codes generated by the fine-tuned model to complete code completion. The input is (x 1 ,x 2 ,…,x t-1 ,y t ,y t+1 ,…,y t+l-1 ), and parallelly output the candidate code sequence S ′ =(y t ′ ,y t ′ +1 ,…,y t ′ +l-1 ).
[0063] 3.3. Using classifier to fuse S and S′: Starting from position t, judge one by one. If the codes output by the large language model and the fine-tuned model at the same position i are different, that is, y i ≠y i ′ , extract features from the output information of the language model when generating the corresponding code and input them into the classifier. The classifier outputs a label. If the label is True, it means that the output y of the fine-tuned model is adopted i , and then judge the next position; if the label is False, it means that the output y of the large language model is adopted i ′ , end the current iteration, and obtain the complete output (y t ,…,y i-1 ,y i ′ ), and start again from step 3.1 until the maximum output length is reached or a complete line of code is completed.
[0064] Example 1
[0065] The present invention evaluates the effectiveness of the proposed method based on a newly constructed domain-specific line-level code completion dataset. This dataset contains 4 mature development domains, where the Django and Flask domains come from the Python community, and the Spring and Android domains come from the Java community. For each domain, the present invention collects high-quality open-source projects with more than 100 stars created after February 2023 from the Github open-source community, extracts the code to construct the dataset, divides the training set to fine-tune the small-scale language model, and divides the test set to construct line-level code completion test examples for evaluation. The present invention uses Deepseek-Coder, one of the most popular current code large models, for evaluation. The small-scale language model uses the GPTQ quantization version of Deepseek-Coder-1.3b, and the large language model uses the Deepseek-Coder-6.7b model. When evaluating the effectiveness of line-level code completion, the present invention uses two commonly used evaluation metrics in the code completion field, EM and ES. The former is used to evaluate whether the completion result of the method is exactly the same as the correct code, and the latter evaluates the edit distance between the completion result of the method and the correct code. The higher the two metrics, the stronger the effectiveness of the method. The present invention compares the proposed method with the language model, the fine-tuned model, and two typical studies in domain-specific code completion, kNM-LM and FT2Ra. Table 1 shows the experimental evaluation results of the method proposed by the present invention on each project dataset.
[0066] Table 1: Evaluation results of the method proposed by the present invention on each project dataset
[0067]
[0068] Three typical open-source projects with the largest number of files were selected from each field for dataset construction and evaluation. The experimental results show that the method of the present invention is significantly superior to the existing methods and the model itself in terms of EM and ES evaluation metrics. In particular, in terms of the EM metric, the method proposed in the present invention has improved by 9.13% compared to the basic large language model and by 7.42% compared to the current optimal research method FT2Ra, demonstrating the effectiveness of the method proposed in the present invention. Table 2 shows the experimental evaluation results of the method proposed in the present invention on the datasets of each field.
[0069] Table 2: Evaluation Results of the Method Proposed in the Present Invention on the Datasets of Each Field
[0070]
[0071] The experiment evaluated each field, and the experimental results show that the method proposed in the present invention is also significantly superior to the existing methods and the model itself in domain-specific code completion. In particular, in terms of the EM metric, the method proposed in the present invention has improved by 5.57% compared to the basic large language model and by 4.67% compared to the current optimal research method FT2Ra, further demonstrating the effectiveness of the method proposed in the present invention. Table 3 shows the evaluation results of the contribution of the features extracted by the method proposed in the present invention to the classifier.
[0072] Table 3: Evaluation Results of the Contribution of the Features Extracted by the Method Proposed in the Present Invention to the Classifier
[0073] Feature The present invention -EMB -PT -TW -LW -LL -FW Accuracy / % 81.1 74.7 80.4 79.7 80.2 80.2 80.5 Feature -FL -SL -SS -RFL -RFS -RSL -RSS Accuracy / % 80.7 80.0 79.9 79.6 79.6 80.6 79.1
[0074] The evaluation method is to sequentially remove the feature data of a certain dimension and evaluate the accuracy of the classifier. The evaluation results show that after removing a certain dimension, the accuracy of the classifier has decreased, indicating that each dimension feature makes a positive contribution to the classifier. In particular, the EMB dimension feature makes the largest contribution to the classifier, corresponding to the fact that the information contained in this feature itself is the most among the features of each dimension.
[0075] After considering the specification and the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary.
[0076] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A domain-specific code completion method based on model collaborative reasoning, characterized in that: The following steps are involved: Collect the code of open source projects in specific fields, build a field-specific code completion dataset, and divide the dataset into a training set, a validation set, and a test set; use the PEFT method to fine-tune the small-scale language model based on the code in the training set, so that the small-scale language model learns the field-specific code knowledge; Extracting codes from the training set and the validation set, constructing line-level code completion samples respectively, and using the large language model and the fine-tuned small-scale language model to reason on the constructed samples respectively to complete the code completion task; During the inference process, the inference results of the large language model and the fine-tuned small-scale language model are collected, and based on the outputs of the large language model and the fine-tuned small-scale language model, multi-dimensional features are retrieved from the input code context and the training set to be completed, so as to train a domain-specific classifier model; Construct a model collaborative reasoning framework and complete reasoning in combination with the inference decoding algorithm. The reasoning process is as follows: first, the fine-tuned small-scale language model performs serial reasoning, and then the large language model generates in parallel based on the reasoning results of the small-scale language model to improve the reasoning speed of the large language model. After the inference is completed, the classifier model is used to select the parts of the output codes of the large language model and the fine-tuned small-scale language model that are different, so as to fuse the inference results of the large language model and the fine-tuned small-scale language model.
2. The domain-specific code completion method based on model collaborative reasoning according to claim 1, characterized in that: The number of parameters of the small-scale language model is within 2B.
3. The domain-specific code completion method based on model collaborative reasoning according to claim 1, characterized in that: The multi-dimensional features include model output dimension, code type dimension, code length dimension, code frequency dimension, code similarity dimension and retrieval information dimension; The model output dimensions include the following features: Input vector of the last layer of the model: During the model prediction process, the vector input to the last layer of each language model; Probability of model output: During the model prediction process, the probability of each language model outputting a specified code; The code type dimension Features include: Code type: the code type output by each language model, including line breaks, punctuation, keywords, identifiers, numbers, and EOS special tokens output by the model; The code length dimension includes the following features: Current completed code length: the length of the completed code of the current model; Input code length: the length of the previous line to be completed given to the model; The code frequency dimension includes the following features: The frequency of occurrence of the current completed code: the frequency of occurrence of all the codes completed by the current model in the context of the code input to the model; The frequency of occurrence of the current completed line: the frequency of occurrence of the line consisting of all the codes completed by the current model in the context of the code input to the model; The code similarity dimension includes the following features: Similarity of the current completed line: the similarity between the line consisting of all the codes completed by the current model and the code above the input to the model; Similarity of the current completed code snippet: The line consisting of all the codes completed by the current model and the context of the line to be completed input to the model form a code snippet, and the similarity between the code snippet and the context is calculated; The retrieval information dimension includes the following features: Frequency of occurrence of the current completed line in the training set: the frequency of occurrence of all lines of code completed by the current model in the code training set; The frequency of occurrence of the current completed code snippet in the training set. The line composed of all the codes completed by the current model and the context above the completed line constitute a code snippet. The frequency of occurrence of the code snippet in the training set is calculated. Similarity between the current completed line and the training set: the similarity between the line consisting of all the codes completed by the current model and the code training set; Similarity between the current completed code snippet and the training set: The code snippet is composed of a line of all the codes completed by the current model and the context above the completed line. The similarity between the code snippet and the training set is calculated.
4. The domain-specific code completion method based on model collaborative reasoning according to claim 1, characterized in that: The construction of the model collaborative reasoning framework and the combination of the inference algorithm to complete the reasoning are as follows: Use a fine-tuned small-scale language model for serial reasoning to complete the code; the input is the code to be completed above c t =(x1,x2,…,x t-1 ), output candidate code sequence S = (y t ,y t+1 ,…,y t+l-1 ), where x1,x2,…,x t-1 With y t ,y t+1 ,…,y t+l-1 represents the token corresponding to the code, and l represents the length of the code output by the fine-tuned small-scale language model; Based on the input code context and the candidate codes generated by the fine-tuned model, the large language model is used for parallel reasoning to complete the code completion. The input is (x1, x2, …, x t-1 ,y t ,y t+1 ,…,y t+l-1 ), output candidate code sequence S ′ =(y t ′ ,y t ′ +1 ,…,y t ′ +l-1 ), where y t ′-y t ′ +l-1 Represents the token corresponding to the output code of the large language model; Use the classifier to fuse S and S′: Start judging one by one from position t. If the codes output by the large language model and the fine-tuned small-scale language model at the same position j are different, y j ≠y j ′ , according to the output information of the fine-tuned small-scale language model when generating the corresponding code, the features are extracted and input into the classifier, which is selected by the classifier to obtain the complete output code sequence (y t ,y j-1 ,y j ′ ).
Citation Information
Patent Citations
High-efficiency large model structured pruning method and device based on parameters
CN117454962A