Method and apparatus for generating target code
By using pre-training and tuning training methods in the generative model, the problem of inefficient writing of test case code is solved, and more efficient test case code generation is achieved.
Patent Information
- Application Number
- CN202210397342.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In the prior art, test case code needs to be written manually and is inefficient.
By determining the code labeling information of the object to be tested, input it into the trained generation model, generating the corresponding output code information, and then determining the target code for testing. The generative model is obtained by pre-training through the first data set and adjustment training through the second data set.
It improves the accuracy of the generated code of the generative model, instead of manually writing test case code, thereby improving the efficiency of writing test case code.
Smart Images

Figure CN114661616B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly, to a method and apparatus for generating target code. Background Art
[0002] Unit test-driven development is a core practice and technology in agile development. Its principle is to write unit test case code before developing functional code, and the test code determines the product code that needs to be written. Currently, unit test code mainly relies on developers to write manually. However, in the most common scenarios such as branch coverage and error coverage, the code topology of unit test code is usually relatively simple, and developers still need to waste too much practice and energy on writing redundant code, resulting in a serious squeeze on the time spent on writing business code and investing in product innovation research and exploration, which is not conducive to business promotion and product innovation.
[0003] In view of the problem that test case code needs to be written manually in related technologies with low efficiency, no effective solution has been proposed yet. Summary of the Invention
[0004] This application provides a method and apparatus for generating target code to solve the problem that test case code needs to be written manually in related technologies with low efficiency.
[0005] According to one aspect of this application, a method for generating target code is provided. The method includes: determining code annotation information of an object under test; inputting the code annotation information into a trained generation model, and the generation model outputs corresponding output code information, where the generation model is pre-trained through a first data set and then adjusted and trained through a second data set. The first data set includes partial code and corresponding complete code, and the second data set includes input code annotation information and corresponding output code information; determining the target code of the object under test according to the output code information, where the target code is used to test the object under test.
[0006] Optionally, inputting the code annotation information into a trained generation model, and the generation model outputs corresponding output code information includes: inputting the code annotation information into a target network of the generation model, and the target network outputs the output code information, where the target network has multiple layers, each layer is provided with a dense decoding module, a connection module is arranged between adjacent dense decoding modules, and the multiple layers of dense decoding modules are connected according to a decoding shortcut mechanism.
[0007] Optionally, inputting the code annotation information into the target network of the generation model, and outputting the output code information by the target network includes: inputting the code annotation information into the first dense decoding module of the first layer, processing the code annotation information by the first dense decoding module to obtain a first processing result, and sending the processing result to the second dense decoding module of the second layer and subsequent multiple connection modules, wherein the first dense decoding module is directly connected to the second dense decoding module; continuously processing the first processing result by the second dense decoding module to obtain a second processing result, and sending the second processing result to the first connection module between the second layer and the third layer and subsequent other connection modules, and processing the second processing result by the first connection module and then sending it to the third dense decoding module of the third layer; continuously processing the second processing result by the third dense decoding module to obtain a third processing result, and sending the second processing result to the second connection module between the third layer and the fourth layer and subsequent other connection modules; processing by subsequent dense decoding modules, and outputting the output code information by the dense decoding module of the last layer and the last connection module.
[0008] Optionally, inputting the code annotation information into the first dense decoding module of the first layer, and processing the code annotation information by the first dense decoding module to obtain a first processing result includes: inputting the code annotation information into the self-attention module of the first dense decoding module, and obtaining a first decoded information after processing by the self-attention module; sending the first decoded information to the first switching regularization module to obtain a second decoded information, wherein there is a residual connection between the self-attention module and the first switching regularization module, and the first switching regularization module is determined by a combination of a layer regularization function and an instance regularization function; sending the second decoded information to the feed-forward module to obtain a third decoded information; inputting the first decoded information, the second decoded information, and the third decoded information into the second switching regularization module, and outputting the first processing result by the second switching regularization module.
[0009] Optionally, determining the code annotation information of the object under test includes: determining the class file of the object under test; extracting the code annotation information from the class file.
[0010] Optionally, before inputting the code annotation information into the trained generation model and outputting the corresponding output code information by the generation model, the method further includes: generating input information in a preset data format according to the code annotation information, wherein the preset data format includes a start flag, an end flag, and the code annotation information; determining the target code of the object under test according to the output code information includes: extracting the target code from the output code information in the preset data format.
[0011] According to another aspect of the present application, a training method for a generation model of target code is provided, including: obtaining a first data set and a second data set, wherein the sources of the first data set and the second data set are different; pre-training a generation model according to the first data set, wherein the first data set includes partial code and corresponding complete code; in the case where the pre-training is completed, adjusting and training the generation model according to the second data set, wherein the second data set includes input code annotation information and corresponding output code information; and when the adjustment training passes the verification, the training is completed.
[0012] Optionally, obtaining the first data set and the second data set includes: collecting first data and second data from different sources; respectively performing cleaning processing on the first data and the second data to obtain the first data set and the second data set, wherein the first data set is divided into a first training set and a first test set, and the second data set is divided into a second training set and a second test set.
[0013] According to another aspect of the present application, a target code generation device is provided. The device includes: a first determination module, configured to determine code annotation information of a measured object; a generation module, configured to input the code annotation information into a trained generation model, and output corresponding output code information by the generation model, wherein the generation model is pre-trained through a first data set and then adjusted and trained through a second data set, the first data set includes partial code and corresponding complete code, and the second data set includes input code annotation information and corresponding output code information; and a second determination module, configured to determine the target code of the measured object according to the output code information, wherein the target code is used to test the measured object.
[0014] According to another aspect of an embodiment of the present invention, a processor is further provided, and the processor is used to run a program, wherein when the program runs, it executes any one of the above target code generation methods.
[0015] According to another aspect of an embodiment of the present invention, an electronic device is further provided, which is characterized by including one or more processors and a memory, and the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement any one of the above target code generation methods.
[0016] Through this application, the following steps are adopted: determining the code annotation information of the object to be measured; inputting the code annotation information into the trained generation model, and the generation model outputs the corresponding output code information. Among them, the generation model is pre-trained through the first data set and then adjusted and trained through the second data set. The first data set includes partial codes and corresponding complete codes, and the second data set includes the input code annotation information and the corresponding output code information; determining the target code of the object to be measured according to the output code information, where the target code is used to test the object to be measured, solving the problem that the test case code needs to be manually written in the related technology with low efficiency. Furthermore, the accuracy of the code generated by the generation model is improved through pre-training and adjusted training, replacing the manual writing of test case codes, and thus improving the writing efficiency of test case codes. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings that form a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0018] Figure 1 is a flowchart of a method for generating a target code provided by an embodiment of this application;
[0019] Figure 2 is a flowchart of the training of a generation model for generating a target code provided by an embodiment of this application;
[0020] Figure 3 is a schematic diagram of a generation model provided by an embodiment of this application;
[0021] Figure 4 is a schematic diagram of a device for generating a target code provided by an embodiment of this application;
[0022] Figure 5 is a schematic diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine the embodiments to detail this application.
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0025] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of this application described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0026] To solve the problem that the test case code in the related technology needs to be manually written with low efficiency, the following methods have emerged in the related technology: CodePro Analytix unit test code generation framework: built on the Eclipse integrated development environment, capable of generating unit test code for a single function, but it mainly generates unit test code through pre-set rules, unable to handle some common test scenarios without return values and without global variable modifications, only supporting the automated generation of simple test cases, and mainly manually writing test cases based on the JUint framework.
[0027] JUnitGenerator unit test automatic code generation: developed based on the Junit unit test framework, generating unit test code according to pre-set templates, only able to generate simple peripheral code, and requiring developers to manually write core test code.
[0028] However, there are still the following problems: The solution based on manual writing takes a long time, and there may be problems that the test cases do not meet the specifications due to the code styles of programmers; the language platform is fixed and cannot be quickly applied to multiple language platforms; the syntax structure of the generated code is fixed and cannot quickly adapt to the update of the unit test case specifications; the later maintenance cost is relatively high; it is necessary to manually define the corresponding code templates; generating test cases based on pre-agreed rules lacks the ability to actively learn and continuously learn from a large amount of code; for closed test scenarios (without return values and without global variables), the success rate is relatively low; the code readability is relatively poor, etc.
[0029] Based on this, the present application hopes to provide a solution that can solve the above technical problems, and its detailed content will be elaborated in the subsequent embodiments.
[0030] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set between the present system and relevant users or institutions. Before obtaining relevant information, a request for acquisition needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, the relevant information is obtained.
[0031] For the convenience of description, some nouns or terms involved in the embodiments of the present application are described as follows:
[0032] Unit Test Driven-Development (UTDD): Before developing functional code, unit test case code is written first, and the test code determines the product code to be written; it is the most commonly used development method in the agile development mode.
[0033] According to an embodiment of the present application, a method for generating target code is provided.
[0034] Figure 1 It is a flowchart of a method for generating target code according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0035] Step S102, determining the code annotation information of the object under test;
[0036] Step S104, inputting the code annotation information into the trained generation model, and the generation model outputs the corresponding output code information. Among them, the generation model is pre-trained through a first data set and then adjusted and trained through a second data set. The first data set includes partial code and corresponding complete code, and the second data set includes the input code annotation information and the corresponding output code information;
[0037] Step S106, determining the target code of the object under test according to the output code information, where the target code is used to test the object under test.
[0038] Through the above steps, determine the code annotation information of the object under test; input the code annotation information into the trained generation model, and the generation model outputs the corresponding output code information. The generation model is pre-trained through the first dataset and then fine-tuned through the second dataset. The first dataset includes partial codes and corresponding complete codes, and the second dataset includes the input code annotation information and the corresponding output code information; determine the target code of the object under test according to the output code information. The target code is used to test the object under test, solving the problem in the related technology that test case codes need to be written manually with low efficiency. Furthermore, it achieves the effect of improving the accuracy of the code generated by the generation model through pre-training and fine-tuning, replacing manual writing of test case codes, and thus improving the writing efficiency of test case codes.
[0039] The above object under test is an object for which test cases need to be created, which can be a piece of development code or functional code. The object under test can be any one of the carrier files of the above-mentioned tested code. By extracting keywords from the object under test, the code annotation information in the object under test can be obtained. The code annotation information is also the annotation of the code, or the rules of the code, the function or task of the code, etc., usually appearing in Chinese.
[0040] The above generation model is pre-trained by the first dataset and then fine-tuned by the second dataset. The above first dataset can be obtained from publicly available code sources on the Internet, including public code libraries, code industry forums, or publicly available corpora or corpus retrieval systems on the Internet, etc. The first dataset is mainly corpus data. Each piece of data in the first dataset is usually composed of partial fields of a piece of code and the complete code. By pre-training the generation model with the first dataset, the generation model can be greatly improved in the ability to restore the entire corpus from partial corpus fields, so that the model can generate the complete required corpus according to partial corpus fields. The above corpus can be understood as a specific paragraph of code.
[0041] After the pre-training is completed, fine-tuning can be carried out through the second dataset. Fine-tuning requires that the pre-trained model can perform better in the unit test code generation task and at the same time meet the relevant code implementation specifications in the actual production scenario. Therefore, the second dataset will be collected from the existing code library of the server of the business system to which the object under test belongs. Each piece of data in the second dataset is usually composed of the input code table annotation information and the code corresponding to the code annotation information. By fine-tuning the pre-trained generation model with the second dataset, the generation model can have a more accurate ability in meeting the actual scenario to which the object under test belongs.
[0042] The generation model is pre-trained and then fine-tuned using the first data set and the second data set in sequence, so that the generation model has more accurate and efficient performance in generating corresponding test case code based on code annotation information. By inputting the code annotation information of the object under test into the generation model, the generation model outputs corresponding output code information, and the target code of the test case corresponding to the object under test is determined based on the output code information.
[0043] It should be noted that the generation model has certain requirements for the format of the input data and the output data. That is, the input data needs to be in the input format required by the model, and the data output by the model is also in the output format corresponding to the model. That is to say, the above output code information is usually not the required target code, and the corresponding target code can be obtained from the output code information through format conversion or data extraction.
[0044] In this embodiment, the first data set and the second data set can be used to improve the accuracy of the code generated by the generation model through pre-training and fine-tuning, replacing manual writing of test case code, thereby improving the efficiency of writing test case code, and solving the problem of low efficiency in manually writing test case code in the related art.
[0045] Optionally, inputting the code annotation information into the trained generation model, and the generation model outputting corresponding output code information includes: inputting the code annotation information into the target network of the generation model, and the target network outputting the output code information. Among them, the target network has multiple layers, each layer is provided with a dense decoding module, a connection module is provided between adjacent dense decoding modules, and the multiple dense decoding modules are connected according to the decoding shortcut mechanism.
[0046] The above generation model includes a target network. The above generation model can be a Transformer model, and the target network can be the decoder Decoder network of the Transformer model. For example Figure 3As shown in the figure, the overall decoder network includes multiple densely connected decoding modules connected in sequence. This densely connected decoding module is different from the decoding module in the standard decoder network. Specifically, in order to enhance the feature expression of the network, the Decoder Block in the standard decoder network (Decoder) is improved, and the Dense Decoder Block is obtained by adding residual connections. At the same time, the SwitchNorm, which combines LayerNorm and InstanceNorm, is used to replace the original LayerNorm in the decoder Decoder. The purpose of adding InstanceNorm is to enhance the regularization of the network for individual samples and improve the attention of the network to the specific features of a certain type of sample, thereby enhancing the overall discriminative power.
[0047] Since the main part of the generation model is the target network, when the code annotation information is input into the trained generation model and the corresponding output code information is output by the generation model, the code annotation information can be input into the target network of the generation model, and the output code information can be output by the target network.
[0048] Optionally, inputting the code annotation information into the target network of the generation model and outputting the output code information by the target network includes: inputting the code annotation information into the first densely connected decoding module of the first layer, processing it by the first densely connected decoding module to obtain a first processing result, and sending the processing result to the second densely connected decoding module of the second layer and subsequent multiple connection modules. The first densely connected decoding module is directly connected to the second densely connected decoding module; the second densely connected decoding module continues to process the first processing result to obtain a second processing result, and sends the second processing result to the first connection module between the second layer and the third layer and subsequent other connection modules. The first connection module processes the second processing result and sends it to the third densely connected decoding module of the third layer; the third densely connected decoding module continues to process the second processing result to obtain a third processing result, and sends the second processing result to the second connection module between the third layer and the fourth layer and subsequent other connection modules; the subsequent densely connected decoding modules are used for processing, and the output code information is output by the densely connected decoding module of the last layer and the last connection module.
[0049] In the target network, multiple layers of dense decoding modules are connected through multiple connection modules. Except for the dense decoding module in the first layer, a connection module is provided behind each layer of the dense decoding module. The connection module includes downsampling, which can solve the problem of feature attenuation. In addition, after each layer of the dense decoding module obtains the processing result, it will send it to the connection module of the subsequent dense decoding module for the downsampling module in the connection module to process, forming a decoding shortcut Dense shortcut mechanism to obtain a densely connected network. Compared with the original network, this network can better alleviate the problem of feature attenuation and has stronger pattern expression ability.
[0050] Optionally, input the code annotation information into the first dense decoding module of the first layer, and the first dense decoding module processes it to obtain the first processing result, including: input the code annotation information into the self-attention module of the first dense decoding module, and obtain the first decoding information after being processed by the self-attention module; send the first decoding information to the first switching regularization module to obtain the second decoding information, where there is a residual connection between the self-attention module and the first switching regularization module, and the first switching regularization module is determined by a combination of a layer regularization function and an instance regularization function; send the second decoding information to the feed-forward module to obtain the third decoding information; input the first decoding information, the second decoding information, and the third decoding information into the second switching regularization module, and the second switching regularization module outputs the first processing result.
[0051] The above-mentioned switching regularization module is composed of a layer regularization function and an instance regularization function, and their combination method is formalized as follows: SwitchNorm(x) = α·LayerNorm(x) + β·InstanceNorm(x), where α and β are the network learnable parameters of the layer regularization function and the instance regularization function respectively, where LayerNorm and InstanceNorm are the layer regularization function and the instance regularization function respectively, and SwitchNorm is the switching regularization method proposed in this embodiment.
[0052] As Figure 3 shown, in Figure 3In part a, the dense decoding module includes a self-attention module (Masked MultiSelf-attention) and two switching regularization modules (Switch Norm). Among them, the first regularization module is directly connected to the self-attention module, and the other is the second regularization module. A feed-forward module is arranged between the first regularization module and the second regularization module. The input of the first regularization module is the input of the self-attention module and the output of the sub-attention module, and the input of the second regularization module is the output of the previous three modules. Through this connection method, the residual connection in the dense decoding module is realized, the attention of the network to the specific features of a certain type of sample is improved, and thus the overall discrimination ability is enhanced.
[0053] Optionally, determining the code annotation information of the object under test includes: determining the class file of the object under test; extracting the code annotation information from the class file.
[0054] The above object under test is an object for which test cases need to be created, which can be a piece of development code or functional code, and the object under test can be any one of the carrier files of the above-mentioned object under test code. By extracting keywords from the object under test, the code annotation information in the object under test can be obtained. The code annotation information is also the annotation of the code, or the rules of the code, the function or task of the code, etc., and usually appears in Chinese.
[0055] Optionally, before inputting the code annotation information into the trained generation model and having the generation model output the corresponding output code information, the method further includes: generating input information in a preset data format according to the code annotation information, where the preset data format includes a start flag, an end flag, and the code annotation information; determining the target code of the object under test according to the output code information includes: extracting the target code from the output code information in the preset data format.
[0056] For example, the input information includes a start flag beg, input data, and an end flag eof. After obtaining the code annotation information, a start flag needs to be added at the very front of the field of the code annotation information, and an end flag needs to be added at the end of the field of the code annotation information before it can be input into the generation model and recognized by the generation model. The output code information output by the model also has a start flag and an end flag, and they need to be removed to obtain the target code.
[0057] Figure 2 It is a flowchart of the training of a generation model for target code provided by an embodiment of the present application. As Figure 2 shown, according to another aspect of the present application, a method for training a generation model for target code is provided, including:
[0058] Step S202: Obtain a first data set and a second data set, where the sources of the first data set and the second data set are different;
[0059] Step S204: Pre-train the generation model according to the first data set, where the first data set includes partial code and corresponding complete code;
[0060] Step S206: When the pre-training is completed, adjust and train the generation model according to the second data set, where the second data set includes input code annotation information and corresponding output code information;
[0061] Step S208: When the adjustment training passes the verification, the training is completed.
[0062] Through the above steps, by obtaining a first data set and a second data set, where the sources of the first data set and the second data set are different; pre-training the generation model according to the first data set, where the first data set includes partial code and corresponding complete code; when the pre-training is completed, adjusting and training the generation model according to the second data set, where the second data set includes input code annotation information and corresponding output code information; when the adjustment training passes the verification, the training is completed, the problem of low accuracy of the test case code generation model in the related art is solved. Furthermore, the effect of improving the accuracy of the code generated by the generation model through pre-training and adjustment training, replacing manual writing of test case code, and thus improving the writing efficiency of test case code is achieved, so as to solve the problem of low efficiency in manually writing test case code in the related art.
[0063] The above generation model is first pre-trained by the first data set and then adjusted and trained by the second data set. The above first data set can be obtained from publicly available code sources on the Internet, including public code libraries, code industry forums, or publicly available corpora or corpus retrieval systems on the Internet, etc. The first data set is mainly corpus data. Each piece of data in the first data set usually consists of partial fields of a piece of code and the complete code. By pre-training the generation model with the first data set, the generation model can be greatly improved in the ability to restore the entire corpus based on partial corpus fields, so that the model can generate the complete required corpus according to partial corpus fields. The above corpus can be understood as a specific paragraph of code.
[0064] After the pre-training is completed, the second data set can be used for adjustment training. The adjustment training requires that the pre-trained model can have better performance in the unit test code generation task, and at the same time needs to meet the relevant code implementation specifications in the actual production scenario. Therefore, the second data set will be collected from the existing code library of the server of the business system to which the object under test belongs. Each data in the second data set is usually composed of the input code table annotation information and the code corresponding to the code annotation information. By adjusting the pre-trained generation model with the second data set, the generation model can have more accurate capabilities in meeting the actual scenario to which the object under test belongs.
[0065] The generation model is pre-trained and adjusted by the first data set and the second data set, so that the generation model has more accurate and efficient performance in generating corresponding test case codes according to the code annotation information. The code annotation information of the object under test is input into the generation model, and the generation model outputs the corresponding output code information, and the target code of the test case corresponding to the object under test is determined according to the output code information.
[0066] Before training the generative model through the first data set and the second data set, the first data set and the second data set may be processed first. Optionally, obtaining the first data set and the second data set includes: collecting the first data and the second data set from different sources; cleaning the first data and the second data set respectively to obtain the first data set and the second data set, wherein the first data set is divided into a first training set and a first test set, and the second data set is divided into a second training set and a second test set.
[0067] Clean the collected first data or second data, remove redundant, incomplete, and grammatically incorrect corpus records, search and replace keyword expressions, and ensure data consistency. In particular, for the first data set used in the model pre-training stage, since the source is public on the Internet, it is necessary to carefully perform cross-dataset weight reduction processing to avoid overfitting the model on certain common corpora. In order to ensure the consistency of keyword expressions (such as Java and java are unified as Java), the FlashText text search and replacement tool can be used to process keywords.
[0068] For the second data set in the model fine-tuning phase, the code has been extracted strictly according to the latest unit test specifications by default during the collection phase, so the corpus is no longer cleaned too much in terms of grammatical specifications. The data set processing work at this stage is mainly on the cleaning of true assertions, invalid assertions and abandoned codes. At the same time, due to considerations of the current algorithm capacity, it is also necessary to remove unit test cases with complex processing logic in the data set, such as the test scenario of using JavaAssist dynamic mapping technology to insert test code.
[0069] It should be noted that this embodiment also provides an alternative implementation, which will be described in detail below.
[0070] This implementation aims to reduce the time consumption of R & D personnel in writing simple test cases, improve work efficiency, and unleash the potential of R & D personnel in product innovation and business focus. At the same time, with intelligent code generation, the writing of unit test scripts will also be more standardized and conducive to maintenance.
[0071] This implementation includes parts such as data processing, model construction, training configuration, verification metrics, and plug-in writing. The implementation solutions and method details of each part will be introduced in detail in turn below.
[0072] Data processing: Data processing can be divided into four parts: data collection, data cleaning, dataset division, and data annotation.
[0073] Data collection: Code generation is essentially still a language generation task, and the training method of the LM (Language Model) model is also applicable to this implementation, that is, model training is divided into two parts: pre-training and fine-tuning. In the pre-training stage, an as general as possible language generation model should be trained, and then in the fine-tuning stage, it should be finely tuned for the unit test code generation task. Therefore, the dataset collection in the pre-training stage should be as applicable as possible to the training of the general language generation model. This implementation uses publicly available corpus dataset resources on the Internet such as Common Crawl dataset, WebText, BookCorpus, English-language Wikipedia, CCL corpus retrieval system, Pre-Modern Chinese Language Corpus, LAMBADA dataset, WebQuestion dataset, and XNLI. In the fine-tuning stage, the pre-trained model is required to perform better in the unit test code generation task and at the same time meet the relevant code implementation specifications in the actual production scenario. Therefore, the dataset will be collected from the existing code library of this center.
[0074] Data cleaning: Before preprocessing the data set, the collected data needs to be cleaned to remove redundant, incomplete, and grammatically incorrect corpus records, find and replace keyword expressions, and ensure data consistency. In particular, for the data set used in the model pre-training stage, it is necessary to carefully reduce the weight of the data set across the data set to avoid overfitting the model on some common corpora. In order to ensure the consistency of keyword expressions (such as Java and java are unified as Java), the FlashText text search and replacement tool is used to process the keywords. For the data set in the model fine-tuning stage, the code has been extracted strictly in accordance with the latest unit test specifications of this center by default during the collection stage, so the corpus is no longer cleaned too much in terms of grammatical specifications. The data set processing work at this stage is mainly on the cleaning of true assertions, invalid assertions, and abandoned codes. At the same time, considering the current algorithm capacity, it is necessary to remove unit test cases with complex processing logic in the data set, such as the test scenario of inserting test code using javaassist dynamic mapping technology.
[0075] Dataset division: The training set in the pre-training stage is a joint dataset consisting of Common Crawl dataset, WebText, BookCorpus, English-language Wikipedia, CCL corpus retrieval system and Pre-Modern ChineseLanguage Corpus, and the test set is LAMBADA dataset, WebQuestion dataset and XNLI, etc. For the dataset in the fine-tuning stage, this implementation adopts the holdout method for division, where 80% of the samples are used for training and 20% of the samples are used for testing.
[0076] Data labeling: In the pre-training stage, this implementation adopts unsupervised training, so no additional data labeling is required; in the fine-tuning stage, a supervised training algorithm is adopted, and the labeling information is the standard unit test case.
[0077] Model construction: This implementation uses the Transformer Decoder network as the benchmark network and proposes a new densely connected network. The network is shown in the figure below. The number of layers and detailed structural parameter configuration of the network are shown in Table 1. Table 1 is a parameter configuration table of the densely connected network:
[0078] Table 1 Parameter configuration table of densely connected networks
[0079] n_layers d_model n_heads d_head 80 10240 80 128
[0080] Figure 3 is a schematic diagram of a generation model provided according to an embodiment of the present application, such as Figure 3As shown, in order to enhance the feature expression of the network, in this embodiment, the standard Decoder Block is improved, and a Dense Decoder Block (also the Dense Block in part a of Figure 3 Figure 3 ) is obtained by adding a residual connection. At the same time, SwitchNorm, which combines LayerNorm and InstanceNorm, is used to replace the LayerNorm in the original Decoder. The purpose of adding InstanceNorm is to enhance the regularization of the network for a single sample and improve the network's attention to the specific features of a certain type of sample, thereby enhancing the overall discriminative ability. The combination method of the two is formalized as follows:
[0081] SwitchNorm(x) = α·LayerNorm(x) + β·InstanceNorm(x)
[0082] where α and β are network-learnable parameters, LayerNorm and InstanceNorm are layer regularization and instance regularization, and SwitchNorm is the switching regularization method proposed in this embodiment.
[0083] Furthermore, in order to alleviate the problem of feature transfer attenuation in deep networks, a Dense shortcut mechanism is proposed for the connection between Dense Decoder Blocks, resulting in a densely connected network. Compared with the original network, this network can better alleviate the problem of feature attenuation and has stronger pattern expression ability.
[0084] Training configuration:
[0085] Warm-up strategy: Cosine warm-up strategy
[0086] Optimizer: Adam, β1 = 0.9, β2 = 0.95, ε = 10 -8
[0087] Weight decay coefficient: 0.1
[0088] Learning rate: 0.7×10 for pre-training -4 , 0.07×10 for fine-tuning -4
[0089] Gradient clipping: Global gradient clipping, clipping interval 1.0
[0090] Objective function:
[0091] Pre-training:
[0092]
[0093] Fine-tuning:
[0094]
[0095] Data input: The data format for pre-training is shown in Table 2, and Table 2 is the data structure table of the first dataset for pre-training:
[0096] Table 2 Data structure table of the first dataset for pre-training
[0097] beg content eof <topred> < / topred> content
[0098] The data format for fine-tuning is shown in Table 3, and Table 3 is the data structure table of the second dataset for fine-tuning:
[0099] Table 3 Data structure table of the second dataset for fine-tuning
[0100] beg content eof <tocode> < / tocode> code
[0101] Among them, <beg>It is the start flag of the input; content is the training sample, which is the corpus sample in the pre-training stage and the class docstring of the unit test code in the fine-tuning stage; <eof>End-of-input marker; <topred>And <tocode>They are all delimiters for training samples and annotation information; content and code are annotation information. During the pre-training stage, the content of content is the same as that of the training samples. During the fine-tuning stage, the content of code is all unit test cases of a single class.
[0102] Since the core algorithm of this embodiment uses an autoregressive model to construct, that is, the prediction of the model is based on the previous prediction of the model, the input and label of the model are almost the same, and the only difference is that the label needs to be shifted one position to the right based on the input.
[0103] Data form:
[0104] Preprocessing stage: No additional processing is done.
[0105] Fine-tuning stage:
[0106] The standard format of docstring is as follows:
[0107] …
[0108] Function description:
[0109] given: Given parameters;
[0110] when: Function call;
[0111] then: Assertion handling;
[0112] …
[0113] The class docstring is the combination of docstrings of all functions in the class under test. Each docstring is separated by <de>Separated by flag symbols.
[0114] For the label data, specifically the file stream of all unit test case codes of the class under test, this file stream will be compressed into a single-line code format, and all spaces between two statements will be replaced by a single space. Similarly, the docstring function description string will be processed in the same way.
[0115] To avoid a large number of potential training biases, spaces will be adopted <sp>Symbol replacement <sp>Will be ignored during prediction and loss calculation.
[0116] Validation metrics: Validation metrics of the pre-trained model:
[0117] ACC (Prediction accuracy):
[0118] Among them, TP is the correctly predicted positive sample, TN is the correctly predicted negative sample, FP is the incorrectly predicted negative sample, and FN is the incorrectly predicted positive sample.
[0119] PPL (Perplexity): Estimate the probability of a sentence appearing based on each word and regularize it with the length of the sentence. The smaller the PPL, the greater the expected probability of the sentence appearing. Therefore, the smaller the PPL, the better. The calculation formula is as follows:
[0120]
[0121] Validation metrics of the fine-tuned model:
[0122] Accuracy:
[0123] Average coverage rate:
[0124] Plugin writing: According to the code language provided by the dataset in the fine-tuning stage, develop corresponding plugin interfaces for the common integrated development environments on that language platform. For example, develop a corresponding plugin for the most popular development environment Idea for the Java language, and integrate the corresponding unit test code generation function in the function right-click menu to achieve the following functions:
[0125] The user right-clicks on the class under test to evoke the unit test code generation function;
[0126] The plugin front-end reads the corresponding class file under test, preprocesses it, and converts it into the input format of the model. Its format is shown in Table 4, and Table 4 is the input data format table of the input model:
[0127] Table 4 Input data format table of the input model
[0128] beg content eof
[0129] After the plugin front-end obtains the model output, it performs post-preprocessing on the result, creates a corresponding test class under the corresponding test path, and writes the generated code;
[0130] A prompt dialog to remind the user whether the generation of unit test cases is completed;
[0131] Provide a feedback interface and a code file upload interface to accept feedback from enthusiastic users, especially code files, for subsequent incremental learning of the model.
[0132] This embodiment introduces a deep LM model into the unit test code generation task; a new densely connected LM model is proposed, which alleviates the feature attenuation problem caused by deep networks and improves the model's expressive ability through dense connection methods within and between Decoder Blocks; a model evaluation metric for evaluating the unit test code generation task is proposed.
[0133] The generation process of unit test cases in this embodiment is completely automated, without the need for additional parameter settings and the writing of rule templates; assuming that the network can converge well, through learning from a large amount of unit test training data, the network can achieve the effect of manually writing unit test cases and is more readable; through learning a large number of unit test scenarios, the network is like an experienced senior language expert and can better give appropriate unit test code for extremely closed scenarios; it can continuously generalize to new unit test scenarios through incremental learning, avoiding cumbersome work such as manually modifying rules and templates; it can be applied to different language platforms through fine-tuning, with lower maintenance costs.
[0134] The method for generating target code provided by an embodiment of this application includes determining the code annotation information of the object under test; inputting the code annotation information into the trained generation model, and the generation model outputs the corresponding output code information, where the generation model is pre-trained through a first data set and then adjusted and trained through a second data set. The first data set includes partial code and the corresponding complete code, and the second data set includes the input code annotation information and the corresponding output code information; determining the target code of the object under test according to the output code information, where the target code is used to test the object under test, solving the problem in the related art that test case code needs to be written manually with low efficiency. Furthermore, it achieves the effect of improving the accuracy of the generated code by the generation model through pre-training and adjusted training, replacing manual writing of test case code, and thus improving the writing efficiency of test case code.
[0135] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0136] An embodiment of this application also provides a device for generating target code. It should be noted that the device for generating target code in an embodiment of this application can be used to execute the method for generating target code provided by an embodiment of this application. The following introduces the device for generating target code provided by an embodiment of this application.
[0137] Figure 4 It is a schematic diagram of a target code generation device according to an embodiment of the present application. As Figure 4 shown, the device includes: a first determination module 42, a generation module 44, and a second determination module 46. The device will be described in detail below.
[0138] The first determination module 42 is configured to determine the code annotation information of the object under test; the generation module 44 is connected to the first determination module 42, and is configured to input the code annotation information into a trained generation model, and the generation model outputs the corresponding output code information. The generation model is pre-trained through a first data set and then adjusted and trained through a second data set. The first data set includes partial codes and corresponding complete codes, and the second data set includes input code annotation information and corresponding output code information; the second determination module 46 is connected to the generation module 44, and is configured to determine the target code of the object under test according to the output code information, where the target code is used to test the object under test.
[0139] Through the above device, the first determination module 42 determines the code annotation information of the object under test; the generation module 44 inputs the code annotation information into a trained generation model, and the generation model outputs the corresponding output code information. The generation model is pre-trained through a first data set and then adjusted and trained through a second data set. The first data set includes partial codes and corresponding complete codes, and the second data set includes input code annotation information and corresponding output code information; the second determination module 46 determines the target code of the object under test according to the output code information, where the target code is used to test the object under test, which solves the problem that test case codes need to be manually written in the related art and the efficiency is low. Furthermore, the effect of improving the accuracy of the code generated by the generation model through pre-training and adjusted training, replacing manual writing of test case codes, and thus improving the writing efficiency of test case codes is achieved.
[0140] Optionally, in the target code generation device provided in the embodiment of the present application, the generation module includes: a generation unit, which also inputs the code annotation information into the target network of the generation model, and the target network outputs the output code information. The target network has multiple layers, each layer is provided with a dense decoding module, a connection module is provided between adjacent dense decoding modules, and the multiple layers of dense decoding modules are connected according to the decoding shortcut mechanism.
[0141] Optionally, the generating unit includes: a first processing subunit, configured to input the code annotation information into the first dense decoding module of the first layer, process the code annotation information by the first dense decoding module to obtain a first processing result, and send the processing result to the second dense decoding module of the second layer and subsequent multiple connection modules, wherein the first dense decoding module is directly connected to the second dense decoding module; a second processing subunit, configured to continue to process the first processing result through the second dense decoding module to obtain a second processing result, and send the second processing result to the first connection module between the second layer and the third layer and subsequent other connection modules, and the first connection module processes the second processing result and sends it to the third dense decoding module of the third layer; a third processing subunit, configured to continue to process the second processing result through the third dense decoding module to obtain a third processing result, and send the second processing result to the second connection module between the third layer and the fourth layer and subsequent other connection modules; a fourth processing subunit, configured to process through subsequent dense decoding modules, and output the output code information by the dense decoding module of the last layer and the last connection module.
[0142] Optionally, the first processing subunit includes: a first decoding secondary subunit, configured to input the code annotation information into the self-attention module of the first dense decoding module, and obtain first decoding information after being processed by the self-attention module; a second decoding secondary subunit, configured to send the first decoding information to the first switching regularization module to obtain second decoding information, wherein there is a residual connection between the self-attention module and the first switching regularization module, and the first switching regularization module is determined by a combination of a layer regularization function and an instance regularization function; a third decoding secondary subunit, configured to send the second decoding information to the feed-forward module to obtain third decoding information; a fourth decoding secondary subunit, configured to input the first decoding information, the second decoding information, and the third decoding information into the second switching regularization module, and output the first processing result by the second switching regularization module.
[0143] Optionally, the first determining module includes: a determining unit, configured to determine the class file of the object under test; an extracting unit, configured to extract the code annotation information from the class file.
[0144] Optionally, the apparatus further includes: a first format conversion module, configured to generate input information in a preset data format according to the code annotation information, wherein the preset data format includes a start flag, an end flag, and the code annotation information; a second format conversion module, configured to determine the target code of the object under test according to the output code information, including: extracting the target code from the output code information in the preset data format.
[0145] The target code generation device provided by the embodiment of the present application, through the first determination module 42, the generation module 44 and the second determination module 46, solves the problem that test case codes need to be manually written in the related art, with low efficiency, and further achieves the effect of improving the accuracy of the generated code by the generation model through pre-training and adjustment training, replacing manual writing of test case codes, and thus improving the writing efficiency of test case codes.
[0146] The above-mentioned target code generation device includes a processor and a memory. The above-mentioned functions of improving the accuracy of the generated code by the generation model through pre-training and adjustment training, replacing manual writing of test case codes, and thus improving the writing efficiency of test case codes are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.
[0147] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the accuracy of the generated code by the generation model can be improved through pre-training and adjustment training, replacing manual writing of test case codes, and thus improving the writing efficiency of test case codes.
[0148] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0149] The embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the target code generation method is implemented.
[0150] The embodiment of the present invention provides a processor, and the processor is used to run a program, wherein when the program runs, the target code generation method is executed.
[0151] Figure 5 is a schematic diagram of an electronic device provided according to the embodiment of the present application, such as Figure 5 As shown in the figure, an embodiment of the present invention provides an electronic device 50, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: determining the code annotation information of the object to be measured; inputting the code annotation information into a trained generation model, and the generation model outputs the corresponding output code information. The generation model is pre-trained through a first data set and then adjusted and trained through a second data set. The first data set includes some codes and the corresponding complete codes, and the second data set includes the input code annotation information and the corresponding output code information; determining the target code of the object to be measured according to the output code information, where the target code is used to test the object to be measured.
[0152] Optionally, inputting the code annotation information into a trained generation model, and the generation model outputs the corresponding output code information, which includes: inputting the code annotation information into the target network of the generation model, and the target network outputs the output code information. The target network has multiple layers, each layer is provided with a dense decoding module, and a connection module is provided between adjacent dense decoding modules. The multiple dense decoding modules are connected according to the decoding shortcut mechanism.
[0153] Optionally, inputting the code annotation information into the target network of the generation model, and the target network outputs the output code information, which includes: inputting the code annotation information into the first dense decoding module of the first layer, and the first dense decoding module processes it to obtain a first processing result, and sends the processing result to the second dense decoding module of the second layer and subsequent multiple connection modules, where the first dense decoding module is directly connected to the second dense decoding module; the second dense decoding module continues to process the first processing result to obtain a second processing result, and sends the second processing result to the first connection module between the second layer and the third layer and subsequent other connection modules, and the first connection module processes the second processing result and sends it to the third dense decoding module of the third layer; the third dense decoding module continues to process the second processing result to obtain a third processing result, and sends the second processing result to the second connection module between the third layer and the fourth layer and subsequent other connection modules; the subsequent dense decoding modules are processed, and the last dense decoding module and the last connection module output the output code information.
[0154] Optionally, input the code annotation information into the first dense decoding module of the first layer, and the first dense decoding module processes it to obtain the first processing result, including: input the code annotation information into the self-attention module of the first dense decoding module, and obtain the first decoded information after being processed by the self-attention module; send the first decoded information to the first switching regularization module to obtain the second decoded information, where there is a residual connection between the self-attention module and the first switching regularization module, and the first switching regularization module is determined by a combination of a layer regularization function and an instance regularization function; send the second decoded information to the feed-forward module to obtain the third decoded information; input the first decoded information, the second decoded information, and the third decoded information into the second switching regularization module, and the second switching regularization module outputs the first processing result.
[0155] Optionally, determining the code annotation information of the object under test includes: determining the class file of the object under test; extracting the code annotation information from the class file.
[0156] Optionally, before inputting the code annotation information into the trained generation model and the generation model outputs the corresponding output code information, the method further includes: generating input information in a preset data format according to the code annotation information, where the preset data format includes a start flag, an end flag, and the code annotation information; determining the target code of the object under test according to the output code information includes: extracting the target code from the output code information in the preset data format.
[0157] When the processor executes the program, the following steps can also be implemented: obtain a first data set and a second data set, where the first data set and the second data set have different sources; pre-train the generation model according to the first data set, where the first data set includes partial codes and corresponding complete codes; in the case of completion of pre-training, adjust and train the generation model according to the second data set, where the second data set includes the input code annotation information and the corresponding output code information; in the case of passing the adjustment training verification, the training is completed.
[0158] Optionally, obtaining the first data set and the second data set includes: collecting the first data and the second data from different sources; respectively performing cleaning processing on the first data and the second data to obtain the first data set and the second data set, where the first data set is divided into a first training set and a first test set, and the second data set is divided into a second training set and a second test set.
[0159] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0160] The present application also provides a computer program product, which when executed on a data processing device, is adapted to execute a program initialized with the following method steps: determining the code annotation information of the object to be measured; inputting the code annotation information into the trained generation model, and outputting the corresponding output code information by the generation model, where the generation model is pre-trained through a first data set and then adjusted and trained through a second data set, the first data set includes partial codes and corresponding complete codes, and the second data set includes the input code annotation information and the corresponding output code information; determining the target code of the object to be measured according to the output code information, where the target code is used to test the object to be measured.
[0161] Optionally, inputting the code annotation information into the trained generation model and outputting the corresponding output code information by the generation model includes: inputting the code annotation information into the target network of the generation model, and outputting the output code information by the target network, where the target network has multiple layers, each layer is provided with a dense decoding module, and a connection module is provided between adjacent dense decoding modules, and the multiple dense decoding modules are connected according to the decoding shortcut mechanism.
[0162] Optionally, inputting the code annotation information into the target network of the generation model and outputting the output code information by the target network includes: inputting the code annotation information into the first dense decoding module of the first layer, processing the code annotation information by the first dense decoding module to obtain a first processing result, and sending the processing result to the second dense decoding module of the second layer and subsequent multiple connection modules, where the first dense decoding module is directly connected to the second dense decoding module; processing the first processing result by the second dense decoding module to obtain a second processing result, and sending the second processing result to the first connection module between the second layer and the third layer and subsequent other connection modules, and processing the second processing result by the first connection module and sending it to the third dense decoding module of the third layer; processing the second processing result by the third dense decoding module to obtain a third processing result, and sending the second processing result to the second connection module between the third layer and the fourth layer and subsequent other connection modules; processing by subsequent dense decoding modules, and outputting the output code information by the dense decoding module of the last layer and the last connection module.
[0163] Optionally, input the code annotation information into the first dense decoding module of the first layer. The first dense decoding module processes it to obtain the first processing result, including: input the code annotation information into the self-attention module of the first dense decoding module, and obtain the first decoding information after being processed by the self-attention module; send the first decoding information to the first switching regularization module to obtain the second decoding information, where there is a residual connection between the self-attention module and the first switching regularization module, and the first switching regularization module is determined by a combination of a layer regularization function and an instance regularization function; send the second decoding information to the feed-forward module to obtain the third decoding information; input the first decoding information, the second decoding information, and the third decoding information into the second switching regularization module, and the second switching regularization module outputs the first processing result.
[0164] Optionally, determining the code annotation information of the object under test includes: determining the class file of the object under test; extracting the code annotation information from the class file.
[0165] Optionally, before inputting the code annotation information into the trained generation model and the generation model outputs the corresponding output code information, the method further includes: generating input information in a preset data format according to the code annotation information, where the preset data format includes a start flag, an end flag, and the code annotation information; determining the target code of the object under test according to the output code information includes: extracting the target code from the output code information in the preset data format.
[0166] The program of the following method steps can also be executed: obtain a first data set and a second data set, where the sources of the first data set and the second data set are different; pre-train the generation model according to the first data set, where the first data set includes partial codes and corresponding complete codes; in the case where the pre-training is completed, adjust and train the generation model according to the second data set, where the second data set includes the input code annotation information and the corresponding output code information; in the case where the adjustment training passes the verification, the training is completed.
[0167] Optionally, obtaining the first data set and the second data set includes: collecting the first data and the second data from different sources; respectively cleaning the first data and the second data to obtain the first data set and the second data set, where the first data set is divided into a first training set and a first test set, and the second data set is divided into a second training set and a second test set.
[0168] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0169] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0172] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0173] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0174] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0175] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0176] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0177] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.< / sp> < / sp> < / de> < / tocode> < / topred> < / eof> < / beg>
Claims
1. A method for generating object code, characterized in that Including: Determine the code annotation information of the object under test; Input the code annotation information into the trained generation model, and the generation model outputs the corresponding output code information. Among them, the generation model is pre-trained through the first data set and then adjusted and trained through the second data set. The first data set includes partial codes and corresponding complete codes, and the second data set includes the input code annotation information and the corresponding output code information; Determine the target code of the object under test according to the output code information, where the target code is used to test the object under test; Among them, inputting the code annotation information into the trained generation model, and the generation model outputs the corresponding output code information includes: inputting the code annotation information into the target network of the generation model, and the target network outputs the output code information. Among them, the target network has multiple layers, each layer is provided with a dense decoding module, and a connection module is arranged between adjacent dense decoding modules. The multiple layers of dense decoding modules are connected according to the decoding shortcut mechanism. The dense decoding module includes a switching regularization module, a feed-forward module, and a self-attention module. The form of the switching regularization module is as follows: SwitchNorm(x) = α·LayerNorm(x) + β·InstanceNorm(x); Where α and β are network learnable parameters, where LayerNorm and InstanceNorm are layer regularization and instance regularization, and SwitchNorm is the switching regularization method; Inputting the code annotation information into the target network of the generation model, and the target network outputs the output code information includes: inputting the code annotation information into the first dense decoding module of the first layer, and the first dense decoding module processes it to obtain a first processing result, and sends the processing result to the second dense decoding module of the second layer, and subsequent multiple connection modules. Among them, the first dense decoding module is directly connected to the second dense decoding module; the second dense decoding module continues to process the first processing result to obtain a second processing result, and sends the second processing result to the first connection module between the second layer and the third layer, and subsequent other connection modules. The first connection module processes the second processing result and sends it to the third dense decoding module of the third layer; the third dense decoding module continues to process the second processing result to obtain a third processing result, and sends the third processing result to the second connection module between the third layer and the fourth layer, and subsequent other connection modules; the subsequent dense decoding modules are processed, and the last dense decoding module and the last connection module output the output code information.
2. The method according to claim 1, characterized in that, Inputting the code annotation information into the first dense decoding module of the first layer, and the first dense decoding module processes it to obtain a first processing result includes: Input the code annotation information into the self-attention module of the first dense decoding module, and obtain the first decoding information after being processed by the self-attention module; Send the first decoding information to the first switching regularization module to obtain the second decoding information. Among them, there is a residual connection between the self-attention module and the first switching regularization module, and the first switching regularization module is determined by a combination of a layer regularization function and an instance regularization function; Send the second decoding information to the feed-forward module to obtain the third decoding information; Input the first decoding information, the second decoding information, and the third decoding information into the second switching regularization module, and output the first processing result by the second switching regularization module.
3. The method according to claim 1, wherein Determining the code annotation information of the object under test includes: Determine the class file of the object under test; Extract the code annotation information from the class file.
4. The method according to claim 3, characterized in that Before inputting the code annotation information into the trained generation model and outputting the corresponding output code information by the generation model, the method further includes: Generate input information in a preset data format according to the code annotation information, where the preset data format includes a start flag, an end flag, and the code annotation information; Determining the target code of the object under test according to the output code information includes: Extract the target code from the output code information in the preset data format.
5. A training method for a target code generation model, characterized in that, Includes: Obtain a first data set and a second data set, where the sources of the first data set and the second data set are different; Pre-train the generation model according to the first data set, where the first data set includes partial codes and corresponding complete codes; When the pre-training is completed, adjust and train the generation model according to the second data set, where the second data set includes input code annotation information and corresponding output code information; When the adjustment training passes the verification, the training is completed; Among them, input the code annotation information into the target network of the generation model, and the target network outputs the output code information. The target network has multiple layers, each layer is provided with a dense decoding module, a connection module is provided between adjacent dense decoding modules, and the multiple dense decoding modules are connected according to the decoding shortcut mechanism; the dense decoding module includes a switching regularization module, a feed-forward module, and a self-attention module, and the form of the switching regularization module is as follows: SwitchNorm(x) = α·LayerNorm(x) + β·InstanceNorm(x); Where α and β are network learnable parameters, where LayerNorm and InstanceNorm are layer regularization and instance regularization, and SwitchNorm is the switching regularization method; Inputting the code annotation information into the target network of the generation model, and outputting the output code information by the target network includes: inputting the code annotation information into the first dense decoding module of the first layer, processing the code annotation information by the first dense decoding module to obtain a first processing result, and sending the processing result to the second dense decoding module of the second layer and subsequent multiple connection modules, where the first dense decoding module is directly connected to the second dense decoding module; processing the first processing result by the second dense decoding module to obtain a second processing result, and sending the second processing result to the first connection module between the second layer and the third layer and subsequent other connection modules, and processing the second processing result by the first connection module and then sending it to the third dense decoding module of the third layer; processing the second processing result by the third dense decoding module to obtain a third processing result, and sending the third processing result to the second connection module between the third layer and the fourth layer and subsequent other connection modules; processing by subsequent dense decoding modules, and outputting the output code information by the dense decoding module of the last layer and the last connection module.
6. The method according to claim 5, wherein Obtaining the first data set and the second data set includes: Collecting the first data and the second data from different sources; Respectively performing cleaning processing on the first data and the second data to obtain the first data set and the second data set, where the first data set is divided into a first training set and a first test set, and the second data set is divided into a second training set and a second test set.
7. An apparatus for generating object code, characterized in that, Including: A first determination module for determining the code annotation information of the object under test; A generation module for inputting the code annotation information into a trained generation model, and outputting corresponding output code information by the generation model, where the generation model is pre-trained by the first data set and then adjusted and trained by the second data set, the first data set includes partial codes and corresponding complete codes, and the second data set includes input code annotation information and corresponding output code information; A second determination module for determining the target code of the object under test according to the output code information, where the target code is used to test the object under test; Among them, the generation module includes: a generation unit for inputting the code annotation information into the target network of the generation model, and outputting the output code information by the target network, where the target network has multiple layers, each layer is provided with a dense decoding module, a connection module is provided between adjacent dense decoding modules, and the multiple layers of dense decoding modules are connected according to the decoding shortcut mechanism; the dense decoding module includes a switching regularization module, a feed-forward module and a self-attention module, and the form of the switching regularization module is as follows: SwitchNorm(x) = α·LayerNorm(x) + β·InstanceNorm(x); where α and β are network learnable parameters, where LayerNorm and InstanceNorm are layer normalization and instance normalization, and SwitchNorm is a switching normalization method; The generating unit includes: a first processing subunit, configured to input the code annotation information into a first dense decoding module of the first layer, process the information by the first dense decoding module to obtain a first processing result, and send the processing result to a second dense decoding module of the second layer and a plurality of subsequent connection modules, wherein the first dense decoding module is directly connected to the second dense decoding module; a second processing subunit, configured to continue processing the first processing result through the second dense decoding module to obtain a second processing result, and send the second processing result to a first connection module between the second layer and the third layer and other subsequent connection modules, and the first connection module processes the second processing result and sends it to a third dense decoding module of the third layer; a third processing subunit, configured to continue processing the second processing result through the third dense decoding module to obtain a third processing result, and send the third processing result to a second connection module between the third layer and the fourth layer and other subsequent connection modules; a fourth processing subunit, configured to perform processing through subsequent dense decoding modules, and output the output code information by a dense decoding module of the last layer and the last connection module.
8. A processor, characterized in that, The processor is configured to run a program, wherein, when the program runs, it executes the method according to any one of claims 1 to 6.
9. An electronic device, characterized in that, Comprising one or more processors and a memory, the memory is configured to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.