Code large model equivalent data enhancement method based on AST abstract syntax tree synonymous replacement
Through a synonymous replacement method based on AST abstract syntax tree, combined with four equivalent replacement methods, syntax semantics checking and data enhancement of the code large model training data is solved, and the problems of low data quality and model overfitting in the existing technology are achieved, efficient and diverse data enhancement is achieved, and model performance is improved.
Patent Information
- Application Number
- CN202510077454.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively and automatically perform syntax semantic checks and data enhancements of code large model training data, resulting in low quality of enhanced data, overfitting of models, and low efficiency of manual recognition and data processing.
A synonymized replacement method based on AST abstract syntax tree is adopted to establish a lexicon of variable names, function names, and class names through data filtering and static syntax analysis, and data enhancement is used to ensure the syntax semantic consistency and diversity of the enhanced data.
It realizes the generation of a large amount of high-quality enhanced data without destroying the code syntax and semantic correctness, which improves the performance and diversity of the code model, reduces the cost of manpower and material resources, and is highly scalable.
Smart Images

Figure CN120010852A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data enhancement methods for intelligent software engineering, and in particular relates to a code large model equivalent data enhancement method based on synonym replacement of AST abstract syntax tree, which can be used to automatically enhance the code large model training corpus to fine-tune and improve the performance of the large language model in vertical fields. Background Art
[0002] Using deep learning and even large language models to complete code-related tasks such as code generation and code completion has always been a hot topic in the field of intelligent software engineering. In recent years, as related technologies have become more mature, large code models have been increasingly widely adopted. The scale and quality of training data are a key factor in determining the performance of large code models. However, on the one hand, the mining of existing open source code data sources (such as Github, Gitee, etc.) has been almost complete, and almost all open source high-quality code data has been used for training, making it difficult to go further; on the other hand, in many specific industry verticals, data within the same industry is often not interoperable for confidentiality reasons. The scale of code training data that a single enterprise can have is relatively limited, and the cost of manually generating new data is high, making it difficult to obtain enough fine-tuning data in this way. At present, a common idea to solve this problem is to regard code as natural language, draw on the existing data enhancement ideas in the field of natural language processing (NLP), and perform synonymous substitution data enhancement on existing code data to expand the scale of the data set. The advantage of this is that while retaining the style and semantics of the original data to the greatest extent, it enhances the diversity of the data, and the required manpower and material costs are relatively low.
[0003] However, there are still several problems that the existing methods have not solved. For example, the existing patents rarely provide an automated solution to ensure the grammatical and semantic validity of the original code data, which sometimes results in low quality of enhanced code data obtained by investing a lot of manpower and material resources; data enhancement is performed only by replacing synonyms and abbreviations, renaming functions and identifiers, etc. (such as patent CN202410264360.6), and the enhanced code data obtained does not have significant diversity, resulting in model overfitting; manual identification of variable names, function names, and class names in the original code data is time-consuming, labor-intensive, and inefficient, especially when facing large-scale data sets, etc. The most important thing is that although the programming language itself has clear grammar and rules, and can be understood as a natural language to a certain extent, the automatic extraction of its hierarchical structure and grammatical framework cannot be directly applied to the existing methods in the NLP field. This requires a low-cost, easy-to-implement, and computer-understandable way to solve the above problems.
[0004] AST abstract syntax tree has come into the attention of researchers because it does not rely on specific grammar, language details, and can be easily modified and adjusted in a code automation way without affecting the original grammatical and semantic correctness. In this case, how to use AST abstract syntax tree to enhance equivalent data with unchanged syntax and semantics based on the characteristics of code data itself has become a new difficulty. Summary of the invention
[0005] The technical problem to be solved by the present invention is to perform synonymous replacement based on the AST abstract syntax tree to automatically construct diversified equivalent training data to improve the performance of the code large model. In order to ensure the grammatical and semantic validity of the enhanced data, data screening and static syntax analysis are first performed, and then the variable names, function names, and class names contained in the code are extracted through the AST abstract syntax tree to establish a vocabulary for screening. On this basis, four equivalent replacement methods are used for data enhancement, and finally it is merged with the original data to obtain the final enhanced code data set.
[0006] The technical solution of the present invention:
[0007] A method for enhancing equivalent data of a large code model based on synonym replacement of an AST abstract syntax tree, the specific steps are as follows:
[0008] Step (1) The initial input is a code data set D consisting of several code files, each of which contains several lines of code. In order to ensure the quality of the original data and reduce the subsequent hardware and time overhead, the code data set D is screened to obtain the screened code data set D fil .
[0009] Step (2) To ensure the grammatical correctness of the input data, filter the code data set D fil All code files in are analyzed using static syntax checking tools. For example, for code files written in Python, the third-party library Pylint is used, for code files written in C++, the cppcheck tool is used, and for Java, the Checkstyle tool is used. For each code file, if its corresponding analysis result contains an error type of "error" or "bug", it is removed from the filtered code dataset D. fil Screen out.
[0010] Step (3) Filter the code data set D filAll code files in the file are extracted using the AST abstract syntax tree to extract the variable names, function names, and class names contained in the code. The corresponding vocabulary of the file is established and filtered, and only variable names, function names, and class names with more than or equal to 2 characters are retained. Then a certain proportion is randomly selected for data enhancement. The proportion can be defined by the user according to the quality and scale of the dataset. All enhanced results are compared with the filtered code dataset D fil Merge and finally obtain the enhanced code dataset D aug .
[0011] Furthermore, step (1) specifically includes the following steps:
[0012] 1-1) Considering that files with too few lines are likely to have no enhancement value, the number of code lines is calculated for all code files in the code dataset D. If the number of code lines corresponding to a file is less than 10, it is screened out from the code dataset D;
[0013] 1-2) Considering that subsequent enhancements are mainly performed on variable names, function names, and class names, if the number is too small, the difference after enhancement will be small, and it will not be significant for fine-tuning the large model. Therefore, for all code files in the code dataset D, AST abstract syntax trees are generated and variable names, function names, and class names in the code are extracted. If the total number of variable names, function names, and class names corresponding to a single code file is less than 20, it will be screened out from the code dataset D.
[0014] Furthermore, step (3) specifically includes the following steps:
[0015] 3-1) For variable names, function names, and class names randomly selected from the vocabulary, first determine whether they belong to camel case naming, underscore naming, or single word naming. For camel case naming and underscore naming types, use predefined code logic to split them into multiple words, and randomly select one of the words with equal probability, and execute step 3-2) for it; for single word naming, directly execute step 3-2) for it;
[0016] 3-2) For each selected word, use a random number to generate a seed, randomly select one of the following synonym replacement methods with equal probability, and perform data enhancement until all variable names, function names, and class names randomly selected from the vocabulary are processed:
[0017] a) Using the industrial-grade natural language processing library Spacy, based on the Spacy text word vector distance generated by convolutional neural network training on the OntoNotes 5 and GloVe Common Crawl datasets, synonym search is performed, the corresponding similarity of the synonyms is calculated, and the synonym with the highest similarity is selected as the pre-selected result. If the similarity of the pre-selected result is greater than or equal to 0.95, it is returned as a replacement result, otherwise it is not replaced and the original word is returned.
[0018] b) Take the first three letters of the selected word as an abbreviation, and return the abbreviated word as the replacement result.
[0019] c) Design a prompt template to be filled: "Please generate prefix and suffix words that match the context for the <variable name / function name / class name content> in the <code file content>". Fill the selected word and the content of the code file in which it is located into the prompt template, obtain the natural language accurate prompt corresponding to the selected word for input into the large language model, call the GPT-3.5TurboAPI interface, and obtain the prefix and suffix word generation results. Set a random number, with a 1 / 3 probability of adding only a prefix to the word, a 1 / 3 probability of adding only a suffix to the word, and a 1 / 3 probability of adding both a prefix and a suffix to the word. The spacing between the added prefix and suffix and the original word should meet the naming format of the variable name / function name / class name where the original word is located, that is, camel case naming, underscore naming, or directly added to the word. Return the added word as the replacement result.
[0020] d) Using the OpenNMT model, the selected word is translated into the corresponding German word and then translated back to English, and the translated word is returned as the replacement result.
[0021] 3-3) For the returned replacement result, determine whether it is composed of multiple words. If so, set a random number seed, and determine whether to adjust the word order with a 50% probability and then proceed to step 3-4); if not, directly proceed to step 3-4);
[0022] 3-4) Use all replacement results to replace the corresponding variable names, function names, and class names in the original code file in turn, and return the replaced code file as the enhancement result.
[0023] Compared with the prior art, the present invention has the following advantages and effects:
[0024] The method of the present invention can efficiently realize the equivalent enhancement of training data for large code models. The present invention utilizes the characteristics of the AST abstract syntax tree that does not rely on specific grammar and language details, and combines a variety of synonymous replacement data enhancement methods to ensure that a large amount of enhanced data is generated without destroying the syntax and semantic correctness of the code, thereby saving a lot of manpower and material resources and improving the performance of large code models; four synonymous replacement methods are randomly selected to ensure the diversity of generated enhanced data and improve the enhancement effect; by accurately judging the naming style of the original variable name, function name, and class name, the style consistency after replacement is ensured, and the deviation that may be introduced by the newly generated data is reduced; in addition, the method of the present invention is also highly scalable and supports a variety of programming languages, which is conducive to improving the user experience and reducing the professional skills required for use. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flow chart of a method for enhancing equivalent data of a large code model based on synonymous replacement of an AST abstract syntax tree according to the present invention.
[0026] Figure 2 It is a process subgraph of the data screening stage in the code large model equivalent data enhancement method based on AST abstract syntax tree synonym replacement of the present invention.
[0027] Figure 3 It is a process subgraph of the data enhancement stage in the code large model equivalent data enhancement method based on AST abstract syntax tree synonym replacement of the present invention. DETAILED DESCRIPTION
[0028] The method of the present invention is described in detail below in conjunction with the accompanying drawings, technical solutions and embodiments.
[0029] like Figure 1 As shown, the code large model equivalent data enhancement method based on AST abstract syntax tree synonym replacement in this embodiment is performed according to the following process: first, the input code data set D consisting of several code files is screened, and files with insufficient lines or too few identifiers are excluded to obtain the screened code data set D fil Then, a static syntax checking tool is used to ensure that the code data set D fil The syntax correctness of the code in the code is checked, and the code files containing syntax errors are filtered out. Next, the D filThe variable names, function names, and class names of the code files are selected and a vocabulary is established. After selecting identifiers with the required number of characters, a certain proportion is randomly selected for data enhancement according to the actual scenario requirements. On the premise of ensuring the consistency of the naming style before and after enhancement, four synonym replacement methods are randomly used for data enhancement, including synonym replacement, abbreviation replacement, large model generation of additional prefixes and suffixes, and translation using the OpenNMT model. Finally, all enhanced results are compared with the filtered code dataset D fil Merge and finally obtain the enhanced code dataset D aug The following uses three Python code files selected from the saurabh0719-elara project of the open source Stack dataset on Hugging Face as an example to explain the implementation details of each process. The specific process is as follows:
[0030] (1) Figure 2 As shown, data screening is performed. The initial input is a code data set D consisting of several code files, each of which contains several lines of code. In order to ensure the quality of the original data and reduce subsequent hardware and time costs, the code data set D is first screened to obtain the screened code data set D fil ;
[0031] 1.1. Considering that files with too few lines are likely to have no enhancement value, the number of code lines is calculated for all code files in the code dataset D. If the number of code lines corresponding to a file is less than 10, it will be screened out from the code dataset D;
[0032] 1.2. Considering that the subsequent enhancement is mainly for variable names, function names, and class names, if the number is too small, the difference after enhancement will be small, and it will not be significant for fine-tuning the large model. Therefore, for all code files in the code dataset D, AST abstract syntax trees are generated and variable names, function names, and class names in the code are extracted. If the total number of variable names, function names, and class names corresponding to a single code file is less than 20, it will be screened out from the code dataset D. Finally, the screened code dataset D is obtained. fil .
[0033] Specifically:
[0034] In this implementation example, the code dataset D includes three Python code files selected from the open source Stack dataset saurabh0719-elara project on Hugging Face. The code file numbers and corresponding programs contained in the code dataset D are as follows:
[0035] Table 1 Code file numbers and corresponding programs contained in code dataset D
[0036]
[0037]
[0038] Next, execute steps 1.1 and 1.2. Calculate the number of code lines and variable names, function names, and class names of the three code files respectively. It is found that the number of code lines corresponding to code file #1 is less than 10, and the number of variable names, function names, and class names contained in code file #87 is less than 20, so they are respectively filtered out from the code dataset D. Finally, the filtered code dataset D is obtained. fil As shown in Table 2 below:
[0039] Table 2 Screening code dataset D fil
[0040]
[0041]
[0042] (2) In order to ensure the grammatical correctness of the input data, the screening code data set D fil All code files in are analyzed using static syntax checking tools. For example, for code files written in Python, the third-party library Pylint is used, for code files written in C++, the cppcheck tool is used, and for Java, the Checkstyle tool is used. For each code file, if its corresponding analysis result contains an error type of "error" or "bug", it is removed from the filtered code dataset D. fil Medium screening;
[0043] Specifically:
[0044] Since the code files used in this example are all written in Python, the third-party library Pylint is used to filter the code dataset D fil The code files contained in the analysis are analyzed, and the corresponding Pylint analysis results are shown in Table 3 below:
[0045] Table 3 Corresponding Pylint analysis results
[0046]
[0047]
[0048] In Pylint, the output results start with a capital letter C, which means the code style is not standard, R means refactoring is recommended, W means warning, and E means error, which means there is a bug in the code. It can be seen that the code file #65 passed the Pylint analysis and does not contain the error type of "error" or "bug", so it is not filtered out.
[0049] (3) Filter code data set D fil All code files in the file are extracted using the AST abstract syntax tree to extract the variable names, function names, and class names contained in the code. The corresponding vocabulary of the file is established and filtered, and only variable names, function names, and class names with more than or equal to 2 characters are retained. Then a certain proportion is randomly selected for data enhancement. The proportion can be defined by the user according to the quality and scale of the dataset. All enhanced results are compared with the filtered code dataset D fil Merge and finally obtain the enhanced code dataset D aug .
[0050] 3.1 For variable names, function names, and class names randomly selected from the vocabulary, first determine whether they belong to camel case naming, underscore naming, or single word naming. For camel case naming and underscore naming types, use predefined code logic to split them into multiple words, and randomly select one of the words with equal probability, and execute step 3.2 for it; for single word naming, directly execute step 3.2 for it;
[0051] 3.2 For each selected word, use a random number to generate a seed, randomly select one of the following synonym replacement methods with equal probability, and perform data augmentation until all variable names, function names, and class names randomly selected from the vocabulary are processed:
[0052] a) Using the industrial-grade natural language processing library Spacy, we perform synonym retrieval based on the Spacy text word vector distance generated by training a convolutional neural network on the OntoNotes 5 and GloVe Common Crawl datasets, calculate the corresponding similarity of the synonyms, and select the synonym with the highest similarity as the pre-selected result. If the pre-selected result similarity is greater than or equal to 0.95, it is returned as a replacement result, otherwise it is not replaced and the original word is returned;
[0053] b) taking the first three letters of the selected word as an abbreviation, and returning the abbreviation as the replacement result;
[0054] c) Design a prompt template to be filled: "Please generate prefixes and suffixes that match the context for the <variable name / function name / class name content> in the <code file content>". Fill the selected word and the content of the code file in which it is located into the prompt template, obtain the natural language accurate prompt corresponding to the selected word for input into the large language model, call the GPT-3.5TurboAPI interface, and obtain the prefix and suffix generation results. Set a random number, with a 1 / 3 probability of adding only a prefix to the word, a 1 / 3 probability of adding only a suffix to the word, and a 1 / 3 probability of adding both a prefix and a suffix to the word. The spacing between the added prefix and suffix and the original word should meet the naming format of the variable name / function name / class name where the original word is located, that is, camel case naming, underscore naming, or directly added to the word. Return the added word as the replacement result;
[0055] d) Using the OpenNMT model, the selected word is translated into the corresponding German word and then translated back to English, and the translated word is returned as the replacement result.
[0056] 3.3 For the returned replacement result, determine whether it is composed of multiple words. If so, set the random number seed, and determine whether to adjust the word order with a 50% probability before entering step 3.4; if not, directly enter step 3.4;
[0057] 3.4 Use all replacement results to replace the corresponding variable names, function names, and class names in the original code file in turn, and return the replaced code file as the enhancement result.
[0058] Specifically:
[0059] First, filter the code dataset D fil The code file contained in the code file, namely code file #65, uses the AST abstract syntax tree to extract the variable names, function names, and class names contained in the code, establishes a vocabulary corresponding to the file and filters it, and only retains the variable names, function names, and class names with more than or equal to 2 characters. The corresponding vocabulary obtained by screening is: "lru_cache, cache_name, cache_size, cache_ttl, keys, delta, cache, persist_exists, clear_cache, add_key, remove_key, cache_info, _ttl, _persist, set_ttl, get_ttl, randomkey, func, total, CacheSystem, LRUCache".
[0060] Then, 30% of them are randomly selected for data augmentation. The proportion can be defined by the user according to the quality and scale of the dataset. Six words are randomly selected from the vocabulary for data augmentation according to the proportion. The randomly selected vocabulary contents are: "keys, delta, randomkey, func, persist_exists, total".
[0061] For randomly selected vocabulary content, first determine whether it belongs to camel case naming, underscore naming, or single word naming, and the corresponding judgment results of the embodiment are shown in Table 4. For camel case naming and underscore naming types, such as persist_exists, the predefined code logic is used to split it into persist and exists, and one of the words persist is randomly selected with equal probability, and single word synonym replacement is performed on it; for other single word naming situations, single word synonym replacement is directly performed on them.
[0062] Table 4 Code file #65 vocabulary naming classification
[0063]
[0064] For single words, including keys, delta, randomkey, func, total, and the single word persist separated by the underscore naming type, a random number is used to generate a seed, and one of the following synonym replacement methods is randomly selected with equal probability to perform data enhancement until all variable names, function names, and class names randomly selected from the vocabulary are processed. The synonym replacement method randomly selected for each variable name, function name, and class name in the vocabulary is shown in Table 5 below:
[0065] Table 5 Code file #65 vocabulary and its synonym replacement method
[0066] Names in the vocabulary Synonym replacement method corresponding to random selection keys c. Add prefixes and suffixes as synonyms delta b. Use abbreviations as synonyms persist b. Use abbreviations as synonyms randomkey d. Translate into a foreign language and then translate back into English as a synonym func c. Add prefixes and suffixes as synonyms total a. Use Spacy, an industrial-grade natural language processing library, to obtain synonyms
[0067] Synonym replacement method:
[0068] a) Using the industrial-grade natural language processing library Spacy, based on the Spacy text word vector distance generated by training the OntoNotes 5 and GloVe Common Crawl datasets using convolutional neural networks, synonym retrieval is performed, the corresponding similarity of the synonyms obtained is calculated, and the synonym with the highest similarity is selected as the pre-selected result. If the similarity of the pre-selected result is greater than or equal to 0.95, it is returned as a replacement result, otherwise no replacement is performed and the original word is returned. Taking the function name total in the vocabulary corresponding to code file #65 as an example, the first five synonyms and the corresponding similarities are shown in Table 6 below:
[0069] Table 6 Total word vector synonyms and their similarities
[0070] Synonyms Similarity Sum 0.98 Entire 0.95 sum 0.95 all 0.93 Whole 0.91
[0071] From Table 6, we can see that the synonym with the closest similarity to total is Sum, so Sum is used as the replacement result to be returned.
[0072] b) Take the first three letters of the selected word as the abbreviation, and return the abbreviated word as the replacement result; taking delta in the vocabulary of code file #65 as an example, select the first three letters of the word as the abbreviation, that is, del, as the replacement result and return it.
[0073] c) Design a prompt template to be filled: "Please generate prefixes and suffixes that match the context for the <variable name / function name / class name content> in the <code file content>". Fill the selected word and the code file content in it into the prompt template, obtain the natural language accurate prompt corresponding to the selected word for input into the large language model, call the GPT-3.5TurboAPI interface, and obtain the prefix and suffix generation results. Set a random number, with a 1 / 3 probability of adding only a prefix to the word, a 1 / 3 probability of adding only a suffix to the word, and a 1 / 3 probability of adding both a prefix and a suffix to the word. The spacing between the added prefix and suffix and the original word should meet the naming format of the variable name / function name / class name where the original word is located, that is, camel case naming, underscore naming, or directly added to the word. Return the added word as the replacement result; taking func in the vocabulary of code file #65 as an example, adding a suffix changes it to func1234 and returns it as the replacement result.
[0074] d) Using the OpenNMT model, the selected word is translated into the corresponding German word and then translated back into English, and the translated word is returned as the replacement result. Taking the randomkey in the vocabulary of code file #65 as an example, it is translated into the German word and then translated back into English Randomly_key, that is, Randomly_key is returned as the replacement result.
[0075] For the returned replacement result, determine whether it is composed of multiple words before replacement. For example, persist_exists is composed of two words before persist is replaced with a synonym. Set a random number seed and determine whether to adjust the word order with a 50% probability. After adjustment, persist_exists eventually becomes exists_per, and then enter the stage of returning the replacement result. If it is not composed of multiple words, directly enter the stage of returning the replacement result.
[0076] In the replacement result return stage, all replacement results are used to replace the corresponding variable names, function names, and class names in the original code file in turn, and the replaced code file is returned as the enhancement result. The result of the replaced code file #65 is as follows:
[0077]
[0078]
[0079] At this point, the data enhancement results are obtained.
Claims
1. A code large model equivalent data enhancement method based on AST abstract syntax tree synonym replacement, characterized in that: The specific steps are as follows: Step (1) The initial input is a code data set D consisting of a number of code files, each of which contains a number of lines of code; the code data set D is filtered to obtain a filtered code data set D fil ; Step (2) Filter the code data set D fil All code files in are analyzed using static syntax checking tools; for each code file, if its corresponding analysis result contains an error type of "error" or "bug", it is removed from the filtered code dataset D fil Medium screening; Step (3) Filter the code data set D fil For all code files in the file, use the AST abstract syntax tree to extract the variable names, function names, and class names contained in the code, build a vocabulary corresponding to the file and filter it, and only keep the variable names, function names, and class names with more than or equal to 2 characters; then randomly select a certain proportion from them for data enhancement, and the proportion can be defined by the user according to the quality and scale of the dataset; Compare all the enhanced results with the filtered code dataset D fil Merge and finally obtain the enhanced code dataset D aug .
2. According to claim 1, a code large model equivalent data enhancement method based on AST abstract syntax tree synonym replacement is characterized in that: Step (1) specifically includes the following steps: 1-1) Calculate the number of code lines for all code files in the code data set D. If the number of code lines corresponding to a file is less than 10, it will be screened out from the code data set D; 1-2) For all code files in the code dataset D, generate AST abstract syntax trees and extract variable names, function names, and class names in the code. If the total number of variable names, function names, and class names corresponding to a single code file is less than 20, it will be screened out from the code dataset D.
3. A method for enhancing the equivalent data of a large code model based on synonymous replacement of an AST abstract syntax tree according to claim 1 or 2, characterized in that: Step (3) specifically includes the following steps: 3-1) For variable names, function names, and class names randomly selected from the vocabulary, first determine whether they belong to camel case naming, underscore naming, or single word naming; for camel case naming and underscore naming types, use predefined code logic to split them into multiple words, and randomly select one of the words with equal probability, and execute step 3-2) for it; for single word naming, directly execute step 3-2) for it; 3-2) For each selected word, use a random number to generate a seed, randomly select one of the following synonym replacement methods with equal probability, and perform data enhancement until all variable names, function names, and class names randomly selected from the vocabulary are processed: a) Using the industrial-grade natural language processing library Spacy, based on the Spacy text word vector distance generated by convolutional neural network training on the OntoNotes 5 and GloVe Common Crawl datasets, synonym retrieval is performed, the corresponding similarity of the synonyms is calculated, and the synonym with the highest similarity is selected as the pre-selected result; if the similarity of the pre-selected result is greater than or equal to 0.95, it is returned as a replacement result, otherwise no replacement is performed and the original word is returned; b) taking the first three letters of the selected word as an abbreviation, and returning the abbreviation as the replacement result; c) Design a prompt template to be filled: "Please generate prefixes and suffixes that match the context for the <variable name / function name / class name content> in the <code file content>"; fill the selected word and the code file content in it into the prompt template, obtain the natural language accurate prompt corresponding to the selected word for input into the large language model, call the GPT-3.5Turbo API interface, and obtain the prefix and suffix generation results; set a random number, with a 1 / 3 probability of adding only a prefix to the word, a 1 / 3 probability of adding only a suffix to the word, and a 1 / 3 probability of adding both a prefix and a suffix to the word; add prefixes and suffixes to the original word in a way that the spacing between them meets the naming format of the variable name / function name / class name where the original word is located, that is, camel case naming, underscore naming, or directly adding to the word; return the added word as the replacement result; d) Using the OpenNMT model, translate the selected word into the corresponding German word and then translate it back to English, and return the translated word as the replacement result; 3-3) For the returned replacement result, determine whether it is composed of multiple words. If so, set a random number seed, and determine whether to adjust the word order with a 50% probability and then proceed to step 3-4); if not, directly proceed to step 3-4); 3-4) Use all replacement results to replace the corresponding variable names, function names, and class names in the original code file in turn, and return the replaced code file as the enhancement result.
Citation Information
Patent Citations
Code annotation generation method and device, electronic equipment and storage medium
CN117850870A
Cited By
Data processing method and device and related equipment
CN121071097A