Equivalent code data enhancement method based on unidirectional translation and validity self-verification

Through the equivalent code data enhancement method based on one-way translation and validity self-verification, large language model and static syntax analysis tools ensure the syntax and semantic effectiveness of candidate codes, the problems of automation and low efficiency of code data enhancement in the existing technology are solved, efficient and accurate code enhancement data generation is achieved, and model performance is improved.

CN120010853APending Publication Date: 2025-05-16DALIAN UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510077737.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to automatically, efficiently and accurately enhancing the code data equivalent, resulting in incorrect syntax, inequality of semantics, and high cost and low operability.

Method used

The equivalent code data augmentation method based on one-way translation and validity self-verification is adopted, and the code functional requirements are translated in one-way through a large language model to generate natural language descriptions. The static syntax analysis tool and the Func2Test model are used to ensure the syntax and semantic validity of candidate codes, and finally the enhanced data set is generated through mixup processing.

Benefits of technology

It realizes automatic, efficient and accurate generation of syntactical and semantic equivalent code-enhanced data, reducing costs, improving operability, and improving the performance of deep learning and large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010853A_ABST
    Figure CN120010853A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of data enhancement methods of intelligent software engineering, and relates to an equivalent code data enhancement method based on one-way translation and validity self-verification. The method comprises the steps that firstly, item content to be subjected to data enhancement is obtained, all contained functions are extracted, function requirement one-way translation is conducted through a large language model in sequence, and natural language description is obtained; and filling a prompt template, and inputting a large language model to generate candidate codes. In order to ensure the grammatical validity of the enhanced data, verifying the candidate code by using a static grammatical analysis tool; in order to ensure semantic equivalence, a Func2Test model is pre-trained in sequence from the perspective of assertion knowledge enhancement and focus method-test case relation learning and fine tuning is carried out, m test cases are generated for each candidate code for testing, and finally only the candidate code with the highest passing rate is reserved as a newly generated code. And finally, performing mix processing on the original code and the newly generated code to obtain a final enhancement result, namely a mixed data set Dmix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data enhancement methods for intelligent software engineering, and in particular, relates to an equivalent code data enhancement method based on one-way translation and self-verification of effectiveness, which can be used to automatically, efficiently and accurately enhance code data for fine-tuning and improving the performance of deep learning and large language models in vertical fields. Background Art

[0002] Deep learning models and large language models that have emerged in recent years have been increasingly widely used in various fields. Since the beginning of the machine learning era, the scale and quality of training data have been regarded as key factors in determining model performance. With the popularity of deep learning methods and pre-trained large language models, the number of parameters of mainstream models has increased exponentially almost every year, and the importance of the scale and quality of training data sets has become increasingly prominent. However, first of all, artificially generated code data is destined to be unable to be used on a large scale due to its extremely high cost, and the scale of its data set is greatly limited; secondly, the public code data on the Internet is often mixed, the quality is uneven, and there are great differences in data distribution. When the amount of data is insufficient, it is easy to cause model underfitting; in addition, many companies cannot open source private data for the sake of protecting commercial secrets, and the internal data style of the data island formed is too single, which may lead to model overfitting. At present, one existing idea to solve this problem is to borrow the two-way translation method in the field of NMT (neural machine translation), translate the code written in a certain programming language into natural language, and then translate it back into another programming language code. The advantage of doing this is that, in theory, the code semantics can be retained as much as possible at a lower cost in manpower and material resources, which enhances data diversity and ensures that data distribution differences are controlled within the original data set.

[0003] However, there are still several unresolved issues with existing methods. For example, the newly generated data cannot guarantee grammatical correctness and semantic equivalence, which makes the time-consuming and laboriously obtained enhanced data unusable or even counterproductive; the existing NMT models are mainly pre-trained based on natural language text rather than code data, and one-way translation cannot be used directly or the effect is poor, let alone two-way translation; in order to verify that the newly generated data is equivalent to the original code data, test cases have to be written manually for verification, and in the end, more manpower and material resources are invested in data set construction, etc. The most important thing is that although the code data can be used to fine-tune the existing NMT model, it faces the logical paradox of making dumplings for a dish of vinegar, which is high cost and low operability. This requires a low-cost, easy-to-implement automated method to solve the above problems.

[0004] Large language models stand out because they have strong understanding and generation capabilities for both code and natural language, are easy to call, and have low time and economic costs. In this case, how to use large language models to enhance code data with grammatically valid and semantically unchanged equivalent data has become a new difficulty. Summary of the invention

[0005] The technical problem to be solved by the present invention is to improve the performance of deep learning and large language models by automatically constructing diversified equivalent training data based on one-way translation and automatic generation of test cases to verify the effectiveness. First, the content of the project to be enhanced is obtained, that is, code, comments, documents, etc. On this basis, the functions contained are extracted, and the functional requirements are translated one-way using the large language model in turn to obtain natural language descriptions; then the prompt template is filled in and input into the large language model to generate n candidate codes for each natural language description. In order to ensure the grammatical validity of the enhanced data, the candidate codes are first verified using a static syntax analysis tool; in order to ensure semantic equivalence, the Func2Test model is pre-trained and fine-tuned from the perspective of assertion knowledge enhancement and focus method-test case relationship learning, and m test cases are generated for each candidate code for testing. Finally, only the candidate code with the highest pass rate is retained as the newly generated code. Finally, the original code and the newly generated code are mixed up to obtain the final enhancement result, that is, the mixed data set D mix .

[0006] The technical solution of the present invention:

[0007] A method for data enhancement of equivalent codes based on one-way translation and validity self-verification, the specific steps are as follows:

[0008] Step (1) Obtain the project content, i.e. the initial data set D, for the project to be enhanced. raw In addition to the code, it may also contain natural language functional descriptions such as code corresponding comments and design documents. Based on the code, corresponding comments, and documents contained in it, a large language model is used to perform a one-way translation of the functional requirements of the functions corresponding to each project content to obtain a natural language functional description dataset D nld .

[0009] Step (2) Construct a one-way translation prompt template for code generation: "Given a function description: <function natural language description>, please use <programming language> to generate the corresponding code." Select any large language model with code generation capability and perform a one-way translation on the dataset D nld For each natural language description in , use the selected model to perform a one-way translation from the natural language description to the corresponding code. For each natural language description, generate n candidate codes for it to form a candidate code dataset D wl .

[0010] Step (3) To ensure the grammatical validity of the candidate code, an open source code grammar analysis tool is used to analyze the candidate code dataset D wl For example, the Pylint tool is used for Python code, the CppCheck tool is used for C++ code, the JavaCompiler is used for Java code, and so on. If the static analysis result of the candidate code contains the "error" or "bug" type, it is removed from the candidate code dataset D wl If all candidate codes fail, re-execute step (2) to obtain a new candidate code dataset D wl Then execute step (3).

[0011] Step (4) To ensure the functional equivalence of the candidate codes, pre-train the Func2Test model and fine-tune it for the candidate code dataset D wl All candidate codes in the code are automatically generated functional test cases to test them and screen candidate codes. Finally, each original function code will get a unique corresponding enhanced code, and the two together form the enhanced code dataset D aug .

[0012] Step (5) Mixup the original code and the newly generated code to generate a new corresponding training data embedding representation and obtain a new mixed data set D mix , this dataset is the final enhanced dataset.

[0013] Furthermore, step (1) specifically includes the following steps:

[0014] 1-1) Construct a one-way translation prompt template for code function requirements: "Given a piece of code written in <programming language>: <code content>, its corresponding comments and document function description are as follows: <function description>, please use <natural language> to describe the function implemented by the code according to the code content";

[0015] 1-2) For the initial data set D raw For each project contained in , identify the programming language used and fill in the <Programming Language> field; and use the AST abstract syntax tree to extract the corresponding function to fill in the <Code Content> field. For each extracted function, design a regular expression according to the project programming language type to extract the comment content contained in it; use the function name as a keyword to search in the design document to extract the corresponding paragraph content. Use the extracted comment content and design document content to fill in the <Function Description> field. Fill in the <Natural Language> field according to the natural language type that the user prefers to use, such as Chinese, English, German, etc. Finally, the initial data set D is obtained. rawA corresponding series of natural language precision prompts for input into a large language model.

[0016] 1-3) Input the natural language precise prompts obtained in step 1-2) into the selected large language model in sequence to generate the initial data set D raw The natural language function descriptions corresponding to each function in all projects in the dataset are summarized to obtain the natural language function description dataset D. nld .

[0017] Furthermore, step (4) specifically includes the following steps:

[0018] 4-1) In order to pre-train the enhanced model assertion knowledge, download the Atlas dataset publicly available in the field and randomly divide it into training set D Atlas-train and validation set D Atlas-val ; In order to pre-train the model to understand the relationship between the focus method and the test case, the Methods2Test dataset publicly available in the field is downloaded and randomly divided into training set D Methods2Test-train and validation set D Methods2Test-val .

[0019] 4-2) Test cases are essentially based on assertions. Therefore, in order to make the Func2Test model have stronger assertion knowledge, it is necessary to pre-train assertion knowledge. Use the PLBART architecture with a masked language model to build a pre-trained assertion language model in a self-supervised manner. Use the [MASK] placeholder to randomly replace the training set D Atlas-train 20% of the tokens in the placeholder are used, and the model is trained to predict the token content corresponding to the placeholder. The ultimate learning goal is to predict the mask token and the corresponding assertion statement of a given focus method. The number of pre-training epochs is set to 100,000, and the model parameters are saved every 10,000 epochs. Atlas-val , use the Accuracy indicator to evaluate the effect, and finally retain the best pre-trained model as the pre-trained Func2Test model. The calculation formula of the Accuracy indicator is shown in formula (1), where TokenNum correct To predict the correct number of tokens, TokenNum all The total number of tokens:

[0020]

[0021] 4-3) Test case generation is essentially strongly related to the model’s understanding of the correspondence between focus methods and test cases, so the Func2Test model is further pre-trained on the correspondence between focus methods and test cases.Methods2Test-train , set the number of pre-training epochs to 10000, save the model parameters every 1000 epochs, and use the validation set D Methods2Test-val , use the ROUGE-L indicator to evaluate the effect, and finally retain the best effect as the final pre-trained Func2Test model. The ROUGE-L indicator is calculated based on the longest common subsequence (LCS), generated and calculated by row, and finally the average value of the ROUGE-L indicator of each row is taken as the total evaluation value. Its calculation formula is shown in formula (2), where LCS ( Generated text, reference text ) That is, the longest common subsequence value between the generated line and the reference text line in the training set, Len 生成文本 That is, generate the line length, Len 参考文本 That is, the reference text line length in the training set:

[0022]

[0023] 4-4) For the candidate code dataset D wl For all candidate codes in the code, use the Func2Test model to generate m corresponding functional test cases and execute them. For each original function code, only the one with the highest proportion of test cases generated by the corresponding n candidate codes is retained.

[0024] 4-5) Finally, each original function code will get a unique corresponding enhanced code, and the two together form the enhanced code dataset D aug .

[0025] Compared with the prior art, the present invention has the following advantages and effects:

[0026] The method of the present invention can efficiently realize the equivalent enhancement of code data for training sets. The present invention utilizes the characteristics of a large language model that has strong understanding and generation capabilities for both code and natural language, and is easy to deploy and has low calling costs. The code function is translated into natural language text and then the corresponding code is regenerated. Combined with the mixup method, it is ensured that semantically equivalent enhanced data with a certain diversity can be generated; by introducing static syntax analysis tools and pre-trained test cases to automatically generate models, the dual validity of generated enhanced data in syntax and semantics is ensured, and the effect of enhanced data on improving model performance is ensured to the greatest extent; in addition, the method of the present invention is also highly scalable, and a variety of large language models can be selected as needed and support a variety of programming languages, which is conducive to improving user experience and reducing the threshold of professional skills required for use. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1It is a flow chart of the equivalent code data enhancement method based on one-way translation and validity self-verification of the present invention.

[0028] Figure 2 It is a flow subgraph of the one-way translation stage in the equivalent code data enhancement method based on one-way translation and validity self-verification of the present invention.

[0029] Figure 3 It is a flow subgraph of the pre-training phase in the equivalent code data enhancement method based on one-way translation and validity self-verification of the present invention. DETAILED DESCRIPTION

[0030] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and the invention content.

[0031] like Figure 1 As shown, the equivalent code data enhancement method based on one-way translation and validity self-verification of the present invention is performed according to the following process: first, the initial data set D consisting of target project code, annotations, and documents is obtained. raw Process it, perform one-way translation of the functional requirements of each function contained in it, and obtain the natural language function description dataset D nld . Then build the code to generate a one-way translation prompt template, using the selected large language model as the dataset D nld Each natural language description generates n candidate codes, forming a candidate code dataset D wl , and use open source code syntax analysis tools for static analysis to filter out data containing bugs or errors. Pre-train the Func2Test model and fine-tune it to create a candidate code dataset D wl All candidate codes in the code are automatically generated functional test cases to test them and screen candidate codes. Finally, each original function code will get a unique corresponding enhanced code, and the two together form the enhanced code dataset D aug The original code and the newly generated code are mixed up to generate a new corresponding training data embedding representation to obtain the final mixed data set D mix The following is an example of a Python code file selected from the saurabh0719-elara project of the open source Stack dataset on Hugging Face to explain the implementation details of each process. The specific implementation methods are as follows:

[0032] (1) For the project to be enhanced, obtain its project content, that is, the initial data set D raw In addition to the code, it may also contain natural language function descriptions such as code corresponding comments and design documents. Based on the code, corresponding comments, and documents contained in it, a large language model is used to perform a one-way translation of the functional requirements of each function to obtain a natural language function description dataset Dnld ;

[0033] 1.1 Construct a one-way translation prompt template for code function requirements: "Given a piece of code written in <programming language>: <code content>, its corresponding comments and document function description are as follows: <function description>, please use <natural language> to describe the function implemented by the code according to its content."

[0034] 1.2 Initial dataset D raw For each project contained in , identify the programming language used and fill in the <Programming Language> field; and use the AST abstract syntax tree to extract the corresponding function to fill in the <Code Content> field. For each extracted function, design a regular expression according to the project programming language type to extract the comment content contained in it; use the function name as a keyword to search in the design document to extract the corresponding paragraph content. Use the extracted comment content and design document content to fill in the <Function Description> field. Fill in the <Natural Language> field according to the natural language type that the user prefers to use, such as Chinese, English, German, etc. Finally, the initial data set D is obtained. raw A corresponding series of natural language precision prompts for input into a large language model.

[0035] 1.3 Input the natural language precise prompts obtained in step 1.2 into the selected large language model in sequence to generate the initial data set D raw The natural language function descriptions corresponding to each function in all projects in the dataset are summarized to obtain the natural language function description dataset D. nld .

[0036] Specifically:

[0037] In this implementation example, the code dataset D includes a Python code file selected from the open source Stack dataset saurabh0719-elara project on Hugging Face. The code file numbers and corresponding programs contained in the code dataset D are shown in Table 1 below:

[0038] Table 1 Code file numbers and corresponding programs contained in code dataset D

[0039]

[0040] Next, execute steps 1.2 and 1.3 to fill in the natural language prompt: "Give a paragraph using <python>Code written:<from typing import List\n def below_zero(operations:List[int])-> bool:\n\t balance=0\n\t for op in operations:\n\t\t balance+=op\n\t\tif balance<0:\n\t\t\t return True\n return False>, the corresponding comments and document function descriptions are as follows: <#Define a function below_zero, which accepts a list of integers as parameters and returns a Boolean value#Parameter: #operations(List[int]): a list of integers, each of which represents an operation on the account balance (positive numbers are deposits, negative numbers are withdrawals)#Return: #bool: If the account balance is lower than zero at any time, True is returned; if the balance is not lower than zero after all operations are completed, False is returned> Please use <Chinese> to describe the functions implemented by the code according to its content", input and select a large language model (here the GPT turbo3.5 model is taken as an example), and finally the natural language function description dataset D is obtained. nld As shown in Table 2 below:

[0041] Table 2 Natural language function description dataset D nld

[0042]

[0043]

[0044] (2) Construct a one-way translation prompt template for code generation: "Given a function description: <function natural language description>, please use <programming language> to generate the corresponding code". Select any large language model with code generation capability and perform a one-way translation on the dataset D nld For each natural language description in , the selected model is used to perform a one-way translation from the natural language description to the corresponding code. For each natural language description, n candidate codes are generated for it, where n is 4, forming a candidate code dataset D wl ;

[0045] Specifically:

[0046] The content of the filled prompt template is: "Given function description: <This code defines a function named below_zero, which accepts an integer list operations as a parameter and returns a Boolean value. The specific implementation and function description of the function are as follows: \nParameters: \n operations(List[int]): This is a list of integers, each integer represents an operation on the account balance (positive numbers represent deposits, negative numbers represent withdrawals). \nReturn value: \n bool: If the account balance is less than zero at any time, it returns True; if the balance is not less than zero after all operations are completed, it returns False. \nImplementation process: \nInitialize a variable balance to record the account balance, with an initial value of 0. \nTraverse each operation op in the operations list: \n①Add the value of the current operation op to balance and update the balance. \n②Check whether the updated balance is less than 0. If so, it returns True immediately. \nIf the account balance is never less than 0 after traversing the entire operations list, it returns False. >, please use <python>Generate corresponding code", input the selected large language model (here the large language model is GPT-4o model as an example), and finally generate the candidate code dataset D consisting of 4 candidate codes wl , the contents of which are shown in Table 3 below:

[0047] Table 3 Candidate code dataset D wl

[0048]

[0049]

[0050] (3) To ensure the grammatical validity of the candidate codes, we use an open source code grammar analysis tool to analyze the candidate code dataset D wl For example, the Pylint tool is used for Python code, the CppCheck tool is used for C++ code, the JavaCompiler is used for Java code, and so on. If the static analysis result of the candidate code contains the "error" or "bug" type, it is removed from the candidate code dataset D wl If all candidate codes fail, re-execute step (2) to obtain a new candidate code dataset D wl Then execute step (3).

[0051] Specifically:

[0052] Since the code files used in this example are written in Python, we use the third-party library Pylint to filter the code dataset D fil The code files contained in the analysis are analyzed, and the corresponding Pylint analysis results are shown in Table 4 below:

[0053] Table 4 Corresponding Pylint analysis results

[0054]

[0055]

[0056] In the analysis results of Pylint, capital letters C at the beginning represent irregular code style; R represents refactoring suggestion; W represents warning; and E represents error, which means there is a bug in the code. It can be seen that the four candidate codes all passed the Pylint analysis, which means that they do not contain "error" or "bug" error types, so they are not screened out.

[0057] (4) To ensure the functional equivalence of the candidate codes, the Func2Test model is pre-trained and fine-tuned for the candidate code dataset D wl All candidate codes in the code are automatically generated functional test cases to test them and screen candidate codes. Finally, each original function code will get a unique corresponding enhanced code, and the two together form the enhanced code dataset D aug .

[0058] 4.1 In order to pre-train the enhanced model assertion knowledge, we download the Atlas dataset publicly available in the field, which is a large dataset containing 188,154 pairs of focus methods and corresponding assertion statements. We randomly divide it into training set D and training set D according to the ratio of 8:2. Atlas-train and validation set D Atlas-val ; In order to enhance the model’s understanding of the relationship between the focus method and the test case, the Methods2Test dataset was downloaded from the public domain, which consists of 780,000 focus methods and corresponding test methods. It was randomly divided into training set D and test set D according to the ratio of 8:2. Methods2Test-train and validation set D Methods2Test-val .

[0059] 4.2 Test cases are essentially based on assertions. Therefore, in order for the Func2Test model to have a stronger basic knowledge of assertions, it is necessary to pre-train assertion knowledge first. Unlike BART, which uses the standard attention mechanism, the PLBART architecture has not only been pre-trained on natural language and 7 programming languages, but also can use the bidirectional attention mechanism to capture context from past and future sequences, thereby effectively learning the input sequence in parallel. Therefore, the PLBART architecture with a masked language model is used to build a pre-trained assertion language model in a self-supervised manner. Use the [MASK] placeholder to randomly replace the training set D Atlas-train 20% of the tokens in the placeholder are used, and the model is trained to predict the token content corresponding to the placeholder. The ultimate learning goal is to predict the mask token and the corresponding assertion statement of a given focus method. The number of pre-training epochs is set to 100,000, and the model parameters are saved every 10,000 epochs. Atlas-val , use the Accuracy indicator to evaluate the effect, and finally retain the best pre-trained model as the pre-trained Func2Test model to enter the next step. The calculation formula of the Accuracy indicator is shown in formula (1), where TokenNum correct To predict the correct number of tokens, TokenNum all The total number of tokens:

[0060]

[0061] 4.3 Test case generation is essentially strongly related to the model’s understanding of the correspondence between focus methods and test cases, so the Func2Test model is further pre-trained on the correspondence between focus methods and test cases. Methods2Test-train , set the number of pre-training epochs to 10000, save the model parameters every 1000 epochs, and use the validation set D Methods2Test-val , use the ROUGE-L indicator to evaluate the effect, and finally retain the best effect as the final pre-trained Func2Test model. The ROUGE-L indicator is calculated based on the longest common subsequence (LCS), generated and calculated by row, and finally the average value of the ROUGE-L indicator of each row is taken as the total evaluation value. Its calculation formula is shown in formula (2), where LCS ( Generating text , Reference text ) That is, the longest common subsequence value between the generated line and the reference text line in the training set, Len 生成文本 That is, generate the line length, Len 参考文本 That is, the reference text line length in the training set:

[0062]

[0063] 4.4 Candidate Code Dataset D wl For all candidate codes in , m corresponding functional test cases are generated and executed respectively by using the Func2Test model, where m is 5. For each original function code, only the one with the highest proportion of corresponding generated test cases among its corresponding n candidate codes is retained. In this embodiment, n is 4.

[0064] 4.5 Finally, each original function code will obtain a unique corresponding enhanced code, and the two together form the enhanced code dataset D aug .

[0065] Specifically:

[0066] Execute steps 4.1, 4.2, 4.3, and 4.4, and finally select the best pre-trained model as the candidate code dataset D wl The generated five corresponding functional test examples are shown in Table 5:

[0067] Table 5 Candidate code dataset D wl Generated 5 corresponding functional test cases

[0068]

[0069]

[0070]

[0071] Execute the above test cases, and finally retain the candidate code content as follows:

[0072] def below_zero(operations:List[int])->bool:

[0073] balance=0

[0074] i=0

[0075] while i <len(operations):

[0076] balance+=operations[i]

[0077] if balance<0:

[0078] return True

[0079] i+=1

[0080] return False

[0081] The two together constitute the enhanced code dataset D aug As shown in Table 6 below:

[0082] Table 6 Enhanced code dataset D aug

[0083]

[0084] (5) Mixup the original code and the newly generated code to generate a new corresponding training data embedding representation and obtain a new mixed data set D mix , this dataset is the final enhanced dataset. The new mixed dataset D mix As shown in Table 7 below:

[0085] Table 7 New mixed dataset D mix

[0086]

[0087] At this point, the data enhancement results are obtained.< / python> < / python>

Claims

1. A method for enhancing equivalent code data based on one-way translation and validity self-verification, characterized in that: The specific steps are as follows: Step (1) Obtain the project content, i.e. the initial data set D, for the project to be enhanced. raw Based on the code, corresponding comments, and documents contained in it, a large language model is used to perform a one-way translation of the functional requirements of the functions corresponding to each project content to obtain a natural language function description dataset D nld ; Step (2) Construct a one-way translation prompt template for code generation: "Given function description: <function natural language description>, please use <programming language> to generate the corresponding code"; select any large language model with code generation capability, and perform a one-way translation on the dataset D nld For each natural language description in , use the selected model to perform a one-way translation from the natural language description to the corresponding code; for each natural language description, generate n candidate codes for it, forming a candidate code dataset D wl ; Step (3) To ensure the grammatical validity of the candidate code, an open source code grammar analysis tool is used to analyze the candidate code dataset D wl All candidate codes in the D are statically analyzed; if the candidate code static analysis results contain "error" or "bug" type, then wl If all candidate codes fail, re-execute step (2) to obtain a new candidate code dataset D wl Then execute step (3); Step (4) To ensure the functional equivalence of the candidate codes, pre-train the Func2Test model and fine-tune it for the candidate code dataset D wl All candidate codes in the code are automatically generated into functional test cases to test them and screen the candidate codes; finally, each original function code will get a unique corresponding enhanced code, and the two together constitute the enhanced code dataset D aug ; Step (5) Mixup the original code and the newly generated code to generate a new corresponding training data embedding representation and obtain a new mixed data set D mix , this dataset is the final enhanced dataset.

2. According to claim 1, the equivalent code data enhancement method based on one-way translation and validity self-verification is characterized in that: Step (1) specifically includes the following steps: 1-1) Construct a one-way translation prompt template for code function requirements: "Given a piece of code written in <programming language>: <code content>, its corresponding comments and document function description are as follows: <function description>, please use <natural language> to describe the function implemented by the code according to the code content"; 1-2) For the initial data set D raw For each project contained in , identify the programming language used and fill in the <Programming Language> field; and use the AST abstract syntax tree to extract the corresponding function to fill in the <Code Content> field; for each extracted function, design a regular expression according to the project programming language type to extract the comment content contained in it; use the function name as a keyword to search in the design document to extract the corresponding paragraph content; use the extracted comment content and design document content to fill in the <Function Description> field; fill in the <Natural Language> field according to the natural language type that the user prefers to use; finally obtain the initial data set D raw A corresponding series of natural language precision prompts for input into a large language model; 1-3) Input the natural language precise prompts obtained in step 1-2) into the selected large language model in sequence to generate the initial data set D raw The natural language function descriptions corresponding to each function in all projects in the dataset are summarized to obtain the natural language function description dataset D. nld .

3. According to claim 1, the equivalent code data enhancement method based on one-way translation and validity self-verification is characterized in that: Step (4) specifically includes the following steps: 4-1) In order to pre-train the enhanced model assertion knowledge, download the Atlas dataset publicly available in the field and randomly divide it into training set D and Atlas-train and validation set D Atlas-val ; In order to pre-train the model to understand the relationship between the focus method and the test case, the Methods2Test dataset publicly available in the field is downloaded and randomly divided into training set D Methods2Test-train and validation set D Methods2Test-val ; 4-2) In order to make the Func2Test model have stronger assertion basic knowledge, it is necessary to pre-train assertion knowledge first; use the PLBART architecture with a masked language model to build a pre-trained assertion language model in a self-supervised manner; use the [MASK] placeholder to randomly replace the training set D Atlas-train 20% of the tokens in the placeholder are used, and the model is trained to predict the token content corresponding to the placeholder. The final learning goal is to predict the mask token and the corresponding assertion statement of a given focus method. The number of pre-training epochs is set to 100,000, and the model parameters are saved every 10,000 epochs. Atlas-val , use the Accuracy index to evaluate the effect, and finally retain the best pre-trained model as the pre-trained Func2Test model; the calculation formula of the Accuracy index is shown in formula (1), where TokenNum correct To predict the correct number of tokens, TokenNum all The total number of tokens: 4-3) Further pre-train the Func2Test model on the correspondence between the focus method and the test case; use the training set D Methods2Test-train , set the number of pre-training epochs to 10000, save the model parameters every 1000 epochs, and use the validation set D Methods2Test-val , the ROUGE-L indicator is used to evaluate the effect, and the best effect is finally retained as the final pre-trained Func2Test model; the ROUGE-L indicator is calculated based on the longest common subsequence LCS, generated and calculated row by row, and finally the average value of the ROUGE-L indicator of each row is taken as the total evaluation value; its calculation formula is shown in formula (2), where LCS ( Generated text, reference text ) That is, the longest common subsequence value between the generated line and the reference text line in the training set, Len 生成文本 That is, generate the line length, Len 参考文本 That is, the reference text line length in the training set: 4-4) For the candidate code dataset D wl For all candidate codes in the code, use the Func2Test model to generate m corresponding functional test cases and execute them in turn; for each original function code, only the one with the highest proportion of corresponding generated test cases among its corresponding n candidate codes is retained; 4-5) Finally, each original function code will get a unique corresponding enhanced code, and the two together form the enhanced code dataset D aug .