Code generation method and related device
By combining search and LLM reasoning, the problem of low efficiency in automated development of test codes in integrated testing and system testing in the existing technology is solved, and the automatic generation of complex test codes is realized, which improves the efficiency of automated development.
Patent Information
- Application Number
- PCT/CN2024/092031
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-05-09
- Publication Date
- 2025-05-22
AI Technical Summary
In the existing technology, in the integration test and system test, the automated development of test code is relatively low and it is difficult to meet business needs.
By combining search and large language model (LLM) inference, text code pairs of vector libraries and test method vector libraries are retrieved to obtain test code snippets and test method call samples corresponding to similar test step descriptions, and to generate test code through LLM inference.
It realizes automatic generation of complex test code, improves the efficiency of automated development of complex test scenarios, and can effectively support integrated testing and system testing.
Smart Images

Figure CN2024092031_22052025_PF_FP_ABST
Abstract
Description
A code generation method and related device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 17, 2023, with application number 202311544947.4 and invention name “A code generation method and related equipment”, and claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 24, 2023, with application number 202311602430.6 and invention name “A code generation method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a code generation method, system, computing device cluster, computer-readable storage medium, and computer program product. Background Art
[0003] From development to release, software typically undergoes a series of tests, including but not limited to unit testing (UT), integration testing, and system testing. Unit testing, integration testing, or system testing can be implemented by executing corresponding test code (e.g., test script code).
[0004] Automated test code development technology for unit testing is relatively mature. With the rise of general-purpose Large Language Models (LLMs), it's now possible to use them to generate function-level unit test code. However, due to the greater complexity of integration or system testing, automated test code development efficiency is low, making it difficult to meet business needs.
[0005] Summary of the Invention
[0006] This application provides a code generation method that combines retrieval and LLM reasoning. By first searching a library of text code pair vectors and a library of test method vectors, the method obtains test code snippets and test method call examples corresponding to similar test step descriptions. Based on these test code snippets and test method call examples, LLM reasoning is performed, rather than reasoning at the function granularity. This method enables the automatic generation of complex test code and improves the efficiency of automated development of complex test scenarios. This application also provides a code generation system, computing device cluster, computer-readable storage medium, and computer program product corresponding to the above method.
[0007] In the first aspect, the present application provides a code generation method. The method can be executed by a code generation system (or referred to as a code generation platform, a code generation tool). The code generation system is used to automatically generate test code. Among them, the code generation system can be a software system, and the software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. It should be noted that the software system can be an independent software package or provide an interface for users to use in the form of a cloud service. The software system can also be integrated into other software for users to use in the form of a plug-in, for example, the software system can be a plug-in for an integrated development environment (IDE). In some examples, the code generation system can also be a hardware system, for example, a computing device cluster with code generation capability, and the computing device cluster executes the code generation method of the present application when it is running.
[0008] Specifically, the code generation system receives the test step description input by the user, and then retrieves the text code pair vector library and the test method vector library based on the test step description to obtain the test code snippets and test method call samples corresponding to the similar test step descriptions, wherein the text code pair vector library includes the description text and the vector formed by the corresponding test code, and the test method vector library includes the test method vector. Then the code generation system generates the test code corresponding to the test step description by reasoning through the large language model (LLM) based on the test code snippets and test method call samples corresponding to the similar test step descriptions. The code generation system presents the test code corresponding to the test step description to the user.
[0009] This method combines retrieval and LLM reasoning. It first searches the text code pair vector library and the test method vector library to obtain test code snippets and test method call examples corresponding to similar test step descriptions (for example, test method call examples with similar semantics). Based on these test code snippets and test method call examples, LLM reasoning is performed instead of reasoning at the function granularity. This can automatically generate complex test code (such as system test code) and improve the efficiency of automated development of complex test scenarios (such as integration testing and system testing).
[0010] In some possible implementations, the code generation system can obtain incremental test text code pairs and incremental test methods from a target domain code repository or integrated development and testing environment. The test method includes at least one of a method definition, a method signature, a parameter description, a method body, and a call example (use case). The code generation system can then vectorize the incremental test text code pairs and add them to a text code pair vector library, and vectorize the incremental test methods and add them to a test method vector library.
[0011] By adopting an incremental update mechanism, the text-code pair vector library can store the latest test step descriptions and test text-code pairs formed by test code, while the test method vector library can store the latest test methods. This ensures the comprehensiveness, accuracy, and timeliness of search results. Inferring and generating test code based on these search results can achieve better results and ensure the quality of the generated test code.
[0012] In some possible implementations, the code generation system can also obtain hints from the context of the test step description. Accordingly, the code generation system can concatenate the hints with test code snippets and test method call examples corresponding to similar test step descriptions. The system then inputs the concatenated hinted test code snippets and test methods into the LLM for inference, obtaining the test code corresponding to the test step description.
[0013] This method can improve the accuracy of reasoning by combining hints in the context, thereby improving the quality of the test code generated by reasoning.
[0014] In some possible implementations, the prompt includes at least one of global information, use case-level information, step-level information, or cross-use case context information in the context. Global information may include product information, file information, and test language framework information. Product information includes product lines and product R&D units; file information includes file names, Git repositories, and file path information; use case-level information includes at least one of test case attribute information, import libraries, use case base class information, and variable initialization definitions; use case attribute information may include at least one of use case class name, use case name, test type, test activity, features, test environment type, preconditions, and test step descriptions; import libraries include at least one of import class signatures, comments, or import method signatures, and comments; use case base class information includes at least one of base class signatures, comments, or base class method signatures, and comments. Step-level information includes at least one of step text descriptions, step test script code blocks, and preceding step codes. Cross-use case context information includes at least one of test suite information, environment information, and configuration information. Test suite information includes at least one of test suite attribute definitions, test suite method signatures, and comments.
[0015] By splicing at least one of the above global information, use case level information, step level information, and cross-use case context information, richer information can be provided to the LLM, thereby improving the accuracy of LLM reasoning in generating test code.
[0016] In some possible implementations, the code generation system can also obtain test engineering code for the target domain and then train a universal base model based on the target domain's test engineering code to obtain an LLM. This allows the trained LLM to automatically generate test code for a specific domain (such as the target domain), rather than being limited to generating test code for general Internet scenarios, thus achieving high usability.
[0017] In some possible implementations, when training a base model to obtain an LLM, the code generation system may pre-train a general base model based on the test engineering code of the target domain to obtain a pre-trained test code generation model. The pre-trained test code generation model may then be fine-tuned in a supervised manner based on text-code pairs extracted from the test engineering code of the target domain to obtain the LLM. Alternatively, the code generation system may fine-tune the general base model in a supervised manner based on text-code pairs extracted from the test engineering code of the target domain to obtain the LLM.
[0018] This method can improve the model training efficiency by pre-training the general base model and then performing supervised fine-tuning on the pre-trained model, or directly performing supervised fine-tuning on the general base model.
[0019] In some possible implementations, the code generation system can also obtain a verification set, which includes text-code pairs in the target domain, and the text-code pairs include descriptive text and real code snippets. The code generation system can then use the LLM model to infer the descriptive text in the verification set to obtain generated code snippets. Accordingly, the code generation system can determine the index value of the evaluation indicator based on the generated code snippet and the real code snippet. Among them, the evaluation indicator includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate or test parameter assignment precision rate. The code generation system can present the index value of the evaluation indicator to the user.
[0020] This method also supports comparing the inference results with the ground truth (GT) and calculating the index values of the evaluation indicators, thereby achieving objective evaluation of LLM and providing a reference for users to decide whether to accept the test code generated by LLM reasoning.
[0021] In some possible implementations, the code generation system can also present the generated code snippet and the real code snippet to the user. This allows the user to directly compare the test code generated by the reasoning with the real test code, thereby enabling a subjective evaluation of the LLM and providing a reference for the user to determine whether to accept the test code generated by the LLM reasoning.
[0022] In some possible implementations, the test step description includes a description based on a natural language or a description based on a domain-specific language, wherein a domain-specific language, also known as a domain-specific language, refers to a computer language that focuses on a certain application domain.
[0023] This method supports generating test code based on natural language or domain-specific language without solidifying the test step description, and has high usability and flexibility.
[0024] In some possible implementations, the code generation system receives a step-level test step description. Accordingly, the code generation system can retrieve a text code pair vector library and a test method vector library based on the step-level test step description, obtain test code snippets and test method call examples corresponding to similar step-level test step descriptions, and then generate step-level test code through LLM reasoning based on the test code snippets and test method call examples. This allows for step-by-step test code generation, allowing for manual intervention or adjustments during the test code generation process to ensure the accuracy of the generated test code.
[0025] In some possible implementations, the code generation system receives a use case-level test step description, which can be a complete test step description. Accordingly, the code generation system can retrieve a text code pair vector library and a test method vector library based on the use case-level test step description, obtain test code snippets and test method call examples corresponding to similar step-level test step descriptions, and generate use case-level test code through LLM reasoning based on the test code snippets and test method call examples.
[0026] This allows test codes for all steps to be generated in batches and returned to the user, improving the efficiency of test code generation.
[0027] In some possible implementations, the code generation system can also perform a specification check on the training corpus. Examples of valid specification check items include the following categories: uniqueness check, validity check, integrity check, consistency check, and security check. Each category can also include several check items. Taking consistency check as an example, consistency check can be subdivided into text code consistency (also known as text script consistency check), use case text quality check, test method quality check, test method density check, and test method style check. Furthermore, text code consistency check can be further subdivided into the following check items: text similarity, inconsistent test code implementation; similar test code implementation, inconsistent text description.
[0028] This method can improve the quality of training data by performing the above-mentioned standard checks and other data quality engineering on the training corpus. Training LLM based on high-quality training data can improve training efficiency and training effects.
[0029] In a second aspect, the present application provides a code generation system. The code generation system is used to automatically generate test code, and the system includes:
[0030] An interactive module, used to receive test step descriptions input by users;
[0031] a retrieval module, configured to retrieve a text-code pair vector library and a test method vector library based on the test step description to obtain test code snippets and test method call examples corresponding to similar test step descriptions, wherein the text-code pair vector library includes vectors formed by description text and corresponding test codes, and the test method vector library includes test method vectors;
[0032] An inference module is configured to generate test code corresponding to the test step description by performing inference using a large language model (LLM) based on the test code snippets and test method call examples corresponding to the similar test step descriptions;
[0033] The interaction module is further configured to present the test code corresponding to the test step description to the user.
[0034] In some possible implementations, the system further includes:
[0035] The vector library management module is used to obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library.
[0036] In some possible implementations, the interaction module is further configured to:
[0037] Get hints from the context of the test step description;
[0038] The reasoning module is specifically used for:
[0039] splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions;
[0040] The test code snippet spliced with the prompt and the test method call sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.
[0041] In some possible implementations, the prompt includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.
[0042] In some possible implementations, the system further includes:
[0043] The training module is used to obtain the test engineering code of the target domain, train the universal base model according to the test engineering code of the target domain, and obtain the LLM.
[0044] In some possible implementations, the training module is specifically configured to:
[0045] Pre-training a universal base model based on the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model based on text-code pairs extracted from the test engineering code of the target domain to obtain the LLM; or
[0046] The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.
[0047] In some possible implementations, the system further includes:
[0048] An evaluation module is configured to obtain a verification set, the verification set including text-code pairs of the target domain, the text-code pairs including descriptive text and real code snippets, use the LLM model to infer the descriptive text in the verification set to obtain generated code snippets, and determine the index value of an evaluation indicator based on the generated code snippets and the real code snippets, the evaluation indicator including at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter value recall rate, and test parameter value precision rate;
[0049] The interaction module is further configured to present the indicator value of the evaluation indicator to the user.
[0050] In some possible implementations, the interaction module is further configured to:
[0051] The generated code snippet and the real code snippet are presented to the user.
[0052] In some possible implementations, the test step description includes a description based on a natural language or a description based on a domain-specific language.
[0053] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, wherein the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, so that the computing device or computing device cluster performs the code generation method described in the first aspect or any implementation of the first aspect.
[0054] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, wherein the instructions instruct a computing device or a computing device cluster to execute the code generation method described in the first aspect or any implementation of the first aspect.
[0055] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or computing device cluster to execute the code generation method described in the first aspect or any one of the implementations of the first aspect.
[0056] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical methods of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments.
[0058] FIG1 is a schematic diagram of a test framework for system testing provided by the present application;
[0059] FIG2 is a schematic diagram of the architecture of a code generation system provided by this application;
[0060] FIG3 is a schematic diagram of a training and reasoning process provided by this application;
[0061] FIG4 is a schematic diagram of a process for performing enhanced reasoning on a vector library and a test method vector library using a retrieval text code provided by the present application;
[0062] FIG5 is a schematic diagram of an evaluation index provided by this application;
[0063] FIG6 is a flowchart of a code generation method provided by the present application;
[0064] FIG7 is a schematic diagram of obtaining a prompt through a prompt project provided by the present application;
[0065] FIG8 is a schematic diagram of a test code corpus inspection specification provided by the present application;
[0066] FIG9 is a flow chart of a data inspection and cleaning tool provided by the present application;
[0067] FIG10 is a schematic diagram of a data inspection and cleaning result provided by the present application;
[0068] FIG11 is a schematic diagram of a framework of a reasoning tool provided by this application;
[0069] FIG12 is a schematic diagram of a result evaluation process provided by this application;
[0070] FIG13 is a schematic diagram of a calculation formula and example of an objective evaluation indicator provided by this application;
[0071] FIG14 is a schematic diagram of the structure of a code generation system provided by the present application;
[0072] FIG15 is a schematic diagram of the structure of a computing device provided by the present application;
[0073] FIG16 is a schematic diagram of the structure of a computing device cluster provided by this application;
[0074] FIG17 is a schematic diagram of the structure of another computing device cluster provided by the present application;
[0075] FIG18 is a schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION
[0076] The terms "first" and "second" in the embodiments of this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.
[0077] First, some technical terms involved in the embodiments of this application are introduced.
[0078] Unit testing, also known as module testing, verifies the correctness of the smallest units of software design (such as program modules or program units). The goal of unit testing is to verify that each program unit correctly implements the functional, performance, interface, and design constraints specified in the detailed design specifications, and to identify potential errors within each unit. Unit testing requires designing test cases based on the program's internal structure. Multiple program units can be independently tested in parallel.
[0079] Integration testing, also known as assembly testing, typically involves sequential, incremental testing of program units (such as software units, components, and subsystems) based on unit testing to check for issues with the interfaces between these units. Integration testing verifies the interface relationships between program units or components, gradually integrating them into program components or software that meet the outline design requirements.
[0080] System testing is to combine the software that has passed the integration test as a part of the system with system elements such as hardware, supporting software, data and platform, and test the quality attributes of the system under test such as system functionality, performance, reliability, security, resilience, maintainability, etc. in a simulated or real environment to discover potential problems and defects in the software, detect code logic and operating results, and verify whether the system meets user needs.
[0081] Compared with unit testing, the testing process of integration testing and system testing is more complicated. The following example uses system testing to illustrate.
[0082] System testing needs to be run in a real hardware environment or simulated hardware environment (for ease of description, it can also be referred to as a system test environment) that integrates many tested software and platform components. Therefore, the test framework of system testing is usually much more complex than that of unit testing. Referring to the schematic diagram of a test framework for system testing shown in Figure 1, the test framework includes a driver layer, a support layer, a business layer, and an application layer. Among them, the driver layer is also called the underlying communication and message parsing layer, which is used to implement communication interaction with the device, command result parsing, and automated script use case execution framework. The support layer is also called the product support layer, which usually includes the class library of the tested product (such as the test object base library obtained by analyzing, designing, and coding each tested system in the product using the test object model) and some reusable common base libraries of the product. The business layer can include the product command line encapsulation layer (command action word, CAW), the business model layer (business action word, BAW), and the test model layer. The application layer includes the test case description layer. The test method is encapsulated layer by layer and finally forms the application layer test code (usually in script form, also called test script code). This test code can run on the system test environment and instantiate test steps and checkpoints by calling BAW / CAW.
[0083] The complexity of the aforementioned testing framework resulted in low efficiency in developing automated system test code, making it difficult to meet the automation requirements of new use cases. Furthermore, testing must maintain existing functionality without breaking it. As functionality increased, the workload for continuous testing increased, necessitating automation to ensure coverage of existing functionality. With the increasing workload of business testing, the business team struggled to secure the manpower required for automation, and simply increasing manpower was not enough to address the automation issue.
[0084] Automatic test code generation based on general LLM is possible, but it is limited to function-level unit testing scenarios. In complex testing scenarios such as system testing, a test case covers hundreds to thousands or even tens of thousands of functions, making it impossible to generate system test code using a fully white-box context.
[0085] In view of this, the present application provides a code generation method. The method can be executed by a code generation system (or referred to as a code generation platform, a code generation tool). The code generation system is used to automatically generate test code. Among them, the code generation system can be a software system, and the software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. It should be noted that the software system can be an independent software package or provide an interface for users to use in the form of a cloud service. The software system can also be integrated into other software for users to use in the form of a plug-in, for example, the software system can be a plug-in for an integrated development environment (IDE). In some examples, the code generation system can also be a hardware system, for example, a computing device cluster with code generation capabilities, and the computing device cluster executes the code generation method of the present application when it is running.
[0086] Specifically, the code generation system receives the test step description input by the user, and then retrieves the text code pair vector library and the test method vector library based on the test step description to obtain the test code snippets and test method call samples corresponding to the similar test step descriptions, wherein the text code pair vector library includes the description text and the vector formed by the corresponding test code, and the test method vector library includes the test method vector. Then the code generation system generates the test code corresponding to the test step description by reasoning through the large language model (LLM) based on the test code snippets and test method call samples corresponding to the similar test step descriptions. The code generation system presents the test code corresponding to the test step description to the user.
[0087] This method combines retrieval and LLM reasoning. It first searches the text code pair vector library and the test method vector library to obtain test code snippets and test method call examples corresponding to similar test step descriptions (for example, test method call examples with similar semantics). Based on these test code snippets and test method call examples, LLM reasoning is performed instead of reasoning at the function granularity. This can automatically generate complex test code (such as system test code) and improve the efficiency of automated development of complex test scenarios (such as integration testing and system testing).
[0088] Moreover, this method can generate test methods for local first-party, second-party, and third-party libraries and test codes for custom test frameworks for business contexts in specific domains through vector libraries in specific domains, such as text code pair vector libraries and test method vector libraries in target domains, without being limited to generating test codes for general scenarios and general test frameworks in the pan-Internet, and has high availability.
[0089] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.
[0090] Referring to the schematic diagram of the architecture of a code generation system shown in FIG2 , the code generation system 20 includes a reasoning subsystem 202, which generates test code through LLM reasoning. Furthermore, the code generation system 20 may also include a data processing subsystem 204, a training subsystem 206, and an evaluation subsystem 208. The data processing subsystem 204 is used to implement data quality engineering (or data engineering), and the training subsystem 206 is used to perform model training based on the data processed by the data processing subsystem 204 to obtain LLM. The reasoning subsystem 202 is used to perform reasoning based on the LLM and generate test code. The evaluation subsystem 208 is used to perform reasoning through the LLM and perform evaluation based on the reasoning results.
[0091] The overall processing flow is described in detail below in conjunction with the system architecture of the code generation system 20.
[0092] In the first phase, data processing subsystem 204 performs data quality engineering. Specifically, data processing subsystem 204 obtains training corpus, processes the training corpus, and obtains training data. Data processing subsystem 204 can be a unified corpus quality inspection and cleaning tool for different data sources. Data processing subsystem 204 can define at least one of a training data format standard or inspection specification. In this embodiment, the training data format standard can include at least one of a test code raw-code corpus data format standard, a test text code pair corpus data format standard, or a test method signature annotation-test code corpus data format standard. Training sets of different training data format standards can be used to train different LLMs. For example, a training set of a test code raw-code corpus data format standard is used to train an L1 LLM. The L1 LLM can be an LLM used to generate test code, obtained by training (e.g., pre-training) a general LLM (general base model, also known as an L0 LLM), also known as a test code generation pre-training model. Based on this, the test code raw-code corpus data format standard can also be referred to as the L1 test code raw-code corpus data format standard. For another example, the test text code pair corpus data format standard or the test method signature annotation-test code corpus data format standard is used to train the L2 LLM. The L2 LLM can be an LLM obtained by directly training the general LLM and used to generate test code, or an LLM obtained by training the above-mentioned L1 LLM and used to generate test code. Among them, supervised fine-tuning (SFT) can be used to train the L2 LLM. Based on this, the above-mentioned test text code pair corpus data format standard and the test method signature annotation-test code corpus data format standard can also be referred to as the L2 test text code pair corpus data format standard and the L2 test method signature annotation-test code corpus data format standard. The inspection specifications can include corpus quality inspection and cleaning specifications, such as L1 / L2 corpus quality inspection and cleaning specifications.
[0093] The data processing subsystem 204 can extract training data from the original test engineering code based on training data format standards. Furthermore, the data processing subsystem 204 can inspect the training data based on inspection specifications. For data that meets the inspection specifications, the data processing subsystem 204 can also cleanse, optimize, and assess data quality characteristics.
[0094] In some possible implementations, the data processing subsystem 204 may also reserve data as a validation set for subsequent evaluation. Based on this, the validation set may also be referred to as an evaluation set.
[0095] In the second phase, the training subsystem 206 performs model training. The training subsystem 206 can obtain domain-specific test engineering code, such as the target domain test engineering code, and train a universal base model based on the target domain test engineering code to obtain an LLM (also called a domain test code LLM, which can be the aforementioned L2 LLM) that can automatically generate target domain test code.
[0096] As shown in Figure 3, the training subsystem 206 selects an L0 LLM (universal base model) suitable for completing text-to-code (e.g., natural language text-to-code) tasks, obtains one or more domain-specific test engineering codes, and trains the L0 LLM using a pre-training method (e.g., incremental pre-training) based on the test engineering codes (e.g., cleaned and optimized test engineering codes) to obtain one or more domain-specific pre-trained models. This pre-trained model is also called a test code generation pre-trained model, which belongs to the L1 LLM. Based on this, the training subsystem 206 extracts text-code pairs from the corresponding one or more domain-specific test engineering codes. This text-code pair can be a test case text-test code snippet pair. The L1 LLM is fine-tuned in a supervised manner to obtain an LLM capable of automatically generating test code. This LLM can generate test code based on domain-specific test text (e.g., test step descriptions), which belongs to the L2 LLM. Optionally, the training subsystem 206 can also directly perform SFT on the L0 LLM to obtain an L2 LLM for subsequent reasoning. Taking the training effect into consideration, the training subsystem 206 may also perform prompt engineering. Accordingly, when performing model training, the prompts obtained from the prompt engineering, such as at least one of the use case global info, dependency library info, and step info prompts, may be combined to perform supervised fine-tuning.
[0097] In the third phase, the reasoning subsystem 202 performs augmented reasoning generation (RAG) based on the retrieved text-code pair vector library and the test method vector library. This is also known as retrieve augmented generation (RAG). The reasoning subsystem 202 can include a front-end plug-in, a back-end intelligent agent (agent), a vector library, and an LLM reasoning service cluster. The vector library includes the aforementioned text-code pair vector library and the test method vector library.
[0098] As shown in Figure 4, during the online reasoning phase, users can input the test step description to be reasoned (which, in some cases, can also include the test case intent) through the front-end plug-in or agent interface. Accordingly, the agent retrieves the text-code pair vector library and the test method vector library to obtain test code snippets corresponding to similar test step descriptions (e.g., the test code snippets corresponding to the top K test step descriptions ranked by similarity) and test method call examples (use cases). The test method call examples can be semantically similar call examples (e.g., the top K call examples ranked by semantic similarity). The agent then combines other context of the test case and, following the prompts from the prompt engineering template instantiation, calls the LLM reasoning service interface (which can be L0 LLM, L1 LLM, or L2 LLM) to obtain the test code generated by LLM reasoning. Furthermore, the agent can perform post-processing on the generated test code, such as truncation, deduplication, and indentation repair. The agent returns the test code to the user through the front-end plug-in or agent interface. The user can provide feedback on the generated test code, such as accepting, modifying, or rejecting it. The reasoning subsystem 202 may store the user's feedback on the generated test code in a feedback database (feedback DB).
[0099] The reasoning subsystem 202 sends a query request to the retrieval retrieval module for a similarity search. After obtaining the search results, it can also determine whether the similarity between the search results and the test step description (e.g., the text to be inferred, including but not limited to the test step description to be inferred) meets the requirements, and decide on subsequent processing based on the judgment result. For example, if the similarity is higher than a set value, the reasoning subsystem 202 can splice the retrieval prompt and perform reasoning through the LLM. For another example, if the similarity is lower than a set value, the reasoning subsystem 202 can discard the retrieval results.
[0100] In some possible implementations, the reasoning subsystem 202 can also be connected to a code repository, such as a domain code repository, a personal test project code repository, or an integrated development and testing environment. Accordingly, in the offline phase, the reasoning subsystem 202 can also obtain incremental test text code pairs and incremental test methods from the code repository of the target domain or the integrated development and testing environment, vectorize the incremental text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library. It should be noted that the above-mentioned incremental data can also be added to the SFT DB for performing SFT on the LLM. In addition, the reasoning subsystem 202 also supports data deletion or aging, including automatic data aging, or deleting data in response to a deletion operation actively triggered by the user.
[0101] It should be noted that the reasoning subsystem 202 can perform reasoning on the validation set before the LLM goes online to evaluate the LLM. This reasoning process is also called model evaluation reasoning. The reasoning subsystem 202 can also perform reasoning after the LLM goes online to evaluate the LLM's operational experience. For example, the reasoning subsystem 202 can generate test code through reasoning and then combine the generated code with the actual code to calculate the code retention results.
[0102] In the fourth stage, evaluation subsystem 208 performs reasoning using the LLM and performs evaluation based on the reasoning results. In some possible implementations, evaluation subsystem 208 may obtain a validation set, which may include text-code pairs in a specific domain (e.g., the target domain). Evaluation subsystem 208 may then use reasoning subsystem 202 to perform objective or subjective evaluation on the reasoning results (code snippets generated based on the description text, referred to as generated code snippets) of the description text in the validation set.
[0103] The objective evaluation process includes: the evaluation subsystem 208 determines the index value of the evaluation indicator based on the generated code snippet and the real code snippet in the verification set. The evaluation indicator includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate, and test parameter assignment precision rate. In some possible implementations, the index value of the evaluation indicator can be calculated using the method shown in Figure 5. The evaluation subsystem 208 can display the index value of the above-mentioned evaluation indicator to the user.
[0104] The subjective evaluation process includes: the evaluation subsystem 208 presents the generated code snippet and the real code snippet to the user, so that the user can compare the generated code snippet and the real code snippet to perform subjective evaluation.
[0105] Based on the aforementioned code generation system 20, the present application provides a code generation method. The code generation method of the present application is described in detail below in conjunction with embodiments.
[0106] Referring to the flowchart of a code generation method shown in FIG6 , the method includes the following steps:
[0107] S602: The code generation system 20 receives a test step description input by a user.
[0108] The test step description is information describing the test steps. In this embodiment, the test step description can be a description based on a natural language (NL) or a description based on a domain-specific language (DSL). Among them, a domain-specific language, also known as a domain-specific language, refers to a computer language that focuses on a certain application field. Different from ordinary cross-domain general-purpose programming languages (GPLs), domain-specific languages are usually used in certain specific fields, such as Hyper Text Markup Language (HTML) for displaying web pages.
[0109] The code generation system 20 supports at least one input method for inputting test step descriptions. For example, the code generation system 20 can present a graphical interface or a command line interface to the user, and the user can input the test step description in text form through the graphical interface or the command line interface. Accordingly, the code generation system 20 can receive the test step description in text form. For another example, the code generation system 20 can provide a voice input control or a voice input switching instruction. When the voice input control or the voice input switching instruction is triggered, the user can input the test step description in voice input form. Accordingly, the code generation system 20 can receive the test step description in voice form. Furthermore, the code generation system 20 can also convert the test step description in voice form into a test step description in text form through voice recognition or speech-to-text technology.
[0110] In the solution of generating test code through LLM reasoning, this embodiment provides multiple scenarios. In one scenario, the user inputs a complete test step description, and the code generation system 20 receives the complete test step description input by the user, so as to generate the test code at one time. In another scenario, the user inputs a partial test step description, and the code generation system 20 receives the partial test step description input by the user, so as to generate the test code step by step. For example, the user can first input the test step description of the first step, and when the code generation system 20 returns the corresponding test code, the user can then input the test step description of the second step and continue to generate the test code for the second step. And so on, which will not be repeated here.
[0111] When inputting the test step description, the user may input the test step description in the form of a comment in the code file. Alternatively, the user may input the test step description in the form of an independent file. This embodiment does not limit this.
[0112] S604: The code generation system 20 searches the text code pair vector library and the test method vector library according to the test step description to obtain test code snippets and test method call examples corresponding to similar test step descriptions.
[0113] The text code pair vector library includes a vector formed by a description text and a corresponding test code, and the test method vector library includes a test method vector. Wherein, the description text refers to the text of the test step description, which can also be called a test text, including but not limited to a test case text. The test method may include at least one of a method definition, a method signature, a parameter description, a method body, and a call example. The above-mentioned vectors in the vector library can be used as an index, for example, the vectors in the text code pair vector library are used to index the corresponding text code pairs, and the vectors in the test method vector library are used to index the corresponding test methods. In order to realize the automatic generation of test codes for specific fields, the text code pair vector library can be a text code pair vector library for a specific field, and the test method vector library can be a test method vector library for a specific field.
[0114] Among them, the text code pair vector library and the test method vector library can adopt an incremental update mechanism to update the database. Specifically, taking a specific field as the target field (such as the financial field) as an example, the code generation system 20 can obtain incremental test text code pairs and incremental test methods from the code warehouse or integrated development test environment of the target field, and then the code generation system 20 can vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library. By adopting the incremental update mechanism, the text code pair vector library can store the latest test step descriptions and test text code pairs formed by the test codes, and the test method vector library can store the latest test methods.
[0115] The code generation system 20 can vectorize the test step description to obtain a test step description vector, and then calculate the distance between the test step description vector and the vector in the text code vector library. The distance can be a cosine distance or a Euclidean distance. The distance between the vectors can be used to characterize the similarity of the vectors. The code generation system 20 can determine similar step descriptions based on the distance between the vectors, and then obtain the test code snippets corresponding to the similar step descriptions. Similarly, the code generation system 20 can perform semantic analysis on the test step description to determine the test method whose semantic similarity meets the requirements. Among them, the semantic similarity meeting the requirements can be that the semantic similarity is greater than a set value. Based on this, the test method is also called a semantically similar test method. The code generation system 20 can obtain semantically similar test method call examples from the test method vector library. The test method call examples corresponding to the test step description can be the above-mentioned semantically similar test method call examples.
[0116] S606 , the code generation system 20 generates test codes corresponding to the test step description by performing reasoning through the LLM based on the test code snippets and test method call examples corresponding to the similar test step descriptions.
[0117] In specific implementation, the code generation system 20 can also obtain hints from the context of the test step description, for example, obtaining hints from the code file where the test step description is located. Accordingly, the code generation system 20 can splice the hints with the test code snippets and test method call samples corresponding to similar test step descriptions, and then input the spliced hint test code snippets and test method call samples into the LLM for reasoning to obtain the test code corresponding to the test step description.
[0118] The prompt can include at least one of global information, use case-level information, step-level information, or cross-use case context information in the context. As shown in Figure 7, global information (referred to as global info) can include product information, file information, and test language framework information. Product information includes product lines and product development units (PDUs). File information includes file names, Git repositories, and file path information. Use case-level information (referred to as use case-level info) includes at least one of TC attribute info, import libraries, use case base class info, and variable initialization definitions. TC attribute info can include at least one of use case class name, use case name, test type, test activity, feature, test environment type, preconditions, and test step description. Import libraries include at least one of import class signatures and comments or import method signatures and comments. Use case base class info includes at least one of base class signatures and comments or base class method signatures and comments. Step-level information (referred to as step-level info) includes at least one of step text description, step test script code block, and preceding step code. Cross-use case context information (referred to as cross-use case info) includes at least one of test suite information (test suite info), environment information (referred to as ENV info), and configuration information (referred to as Config info). Test suite info includes at least one of test suite attribute definitions, test suite method signatures, and annotations.
[0119] An LLM is a model used to automatically generate test code. This model can be a general LLM, such as the L0 LLM, or an LLM trained using training data from the target domain, such as the aforementioned L1 LLM or L2 LLM. L1 LLMs and L2 LLMs can be trained from a general base model, such as the L1 LLM. The LLM training process is described in detail below.
[0120] The code generation system 20 can obtain test engineering code for the target domain and then train a universal base model based on the test engineering code to obtain an LLM. This LLM can automatically generate test code for the target domain. The code generation system 20 can obtain prompts from the context of the test engineering code, use actual code snippets corresponding to test step descriptions in the test engineering code for the target domain as answers, construct prompt-answer pairs based on the prompt and answer, and fine-tune the L0 LLM or L1 LLM through supervised fine-tuning of the prompt-answer pairs to obtain the L2 LLM.
[0121] In some possible implementations, the code generation system 20 can pre-train a general base model (such as the L0 LLM described above) based on the test engineering code of the target domain to obtain a pre-trained test code generation model (such as the L1 LLM described above). The pre-trained test code generation model can be a pre-trained model for the target domain for automatically generating test code. The code generation system 20 can then perform supervised fine-tuning on the pre-trained test code generation model based on text-code pairs extracted from the test engineering code of the target domain to obtain an LLM (such as the L2 LLM described above).
[0122] In some other possible implementations, the code generation system 20 may directly perform supervised fine-tuning on a general base model (such as L0 LLM) based on text-code pairs extracted from test engineering codes in the target domain to obtain an LLM (such as L2 LLM).
[0123] S608 : The code generation system 20 presents the test code corresponding to the test step description to the user.
[0124] Specifically, the code generation system 20 can present the test code corresponding to the test step description to the user through a graphical user interface or a command line interface. Furthermore, the graphical user interface can also include feedback controls, such as an accept control, a modify control, and a reject control. When the user triggers the accept control, it indicates acceptance of the test code generated by LLM reasoning. When the user triggers the modify control, it allows modification of the test code generated by LLM reasoning. When the user triggers the reject control, it indicates rejection of the test code generated by LLM reasoning.
[0125] In some possible implementations, the code generation system 20 may determine the code retention rate and present the code retention rate to the user for objective evaluation. The code generation system 20 may also present the test code generated by LLM reasoning and the user-modified test code to the user for subjective evaluation.
[0126] Based on the above description, the code generation method of the present application combines retrieval and LLM reasoning. By first searching the text code pair vector library and the test method vector library, test code snippets and test method call samples corresponding to similar test step descriptions (for example, test method call samples with similar semantics) are obtained. Based on the test code snippets and test method call samples, LLM reasoning is performed instead of reasoning at the function granularity. In this way, complex test code (such as system test code) can be automatically generated, thereby improving the efficiency of automated development of complex test scenarios (such as integration testing and system testing).
[0127] The above describes the code generation method from the perspective of the code generation system 20. The following describes the code generation method of the present application in detail in conjunction with specific application scenarios.
[0128] In this application scenario, the code generation method may include the following steps:
[0129] Step 1: Standardization, specification checking and cleaning of LLM training corpus data format
[0130] 1. The code generation system 20 determines the standard format of L1 & L2 LLM training corpus data.
[0131] In some examples, the L2 test text code pairs corpus data format is as follows:
[0132] 2. The code generation system 20 uses inspection rules and cleaning tools to check whether the above corpus complies with the specifications according to the major categories, minor categories, and inspection items, and automatically or manually cleans and rectify the violations.
[0133] As shown in Figure 8, for the task of generating script code from system test case text, valid examples of standard check items include the following categories: uniqueness check, validity check, completeness check, consistency check, and security check. Each category can also include several check items. Taking consistency check as an example, consistency check can be subdivided into text code consistency (also known as text script consistency check), use case text quality check, test method quality check, test method density check, and test method style check. Furthermore, text code consistency check can be further subdivided into the following check items: text similarity, inconsistent test code implementation; similar test code implementation, inconsistent text description.
[0134] Figure 9 also shows the process of the data inspection and cleaning tool, and Figure 10 shows the results of the inspection and cleaning according to the above inspection specifications. Figure 10 also provides a sample data quality measurement dashboard.
[0135] In step 2, the code generation system 20 executes the prompt project and performs model training.
[0136] 1) The code generation system 20 incrementally pre-trains L0 based on the test engineering code corpus data after cleaning in step 1 to obtain L1 LLM.
[0137] It should be noted that this step may not be performed when executing the code generation method of the present application. For example, the code generation system 20 may directly train L0 to obtain L2 LLM.
[0138] 2) Prompt Project: Based on the cleaned test project code corpus data from step 1, the code generation system 20 extracts the context of the test case and constructs a prompt. The information used to construct the prompt includes, but is not limited to, the information dimensions shown in Figure 7. For example, the prompt may include global information such as product, path, language, and test framework; use case-level information such as the test case name, purpose, category, full step description, import dependency libraries, and base class information; step-level information such as the text description of the step to be inferred and the code referenced by the inference step; and information such as cross-use case test suites and environment configuration.
[0139] 3) The code generation system 20 constructs prompt-answer pairs based on the real code snippets corresponding to the test step descriptions as answers, and then fine-tunes the L0 LLM or L1 LLM in a supervised manner to obtain the L2 LLM.
[0140] Step 3: The code generation system 20 generates test code by reasoning based on the newly input test step description (e.g., test case text) and the prompts in the context. The framework of a feasible reasoning tool (such as the reasoning subsystem 202) is shown in Figure 11. The reasoning subsystem 202 achieves reasoning enhancement by retrieving the latest local text code vector library and test method vector library. Among them, the reasoning subsystem 202 includes a front-end plug-in and an API interface for interacting with the user to obtain the test step description and context. The task request information including the test step description and context is routed to the corresponding agent service after flow control and authentication through the Agent scheduling system, completing context extraction and vector library similarity information retrieval, and then splicing prompts to request the LLM reasoning service to generate test code by reasoning. The test code can be post-processed and then returned to the API interface through the action module, and finally returned to the user for confirmation through the IDE plug-in.
[0141] It should be noted that the code generation system 20 may include multiple reasoning scenarios when reasoning test codes, which are described below respectively.
[0142] Reasoning scenario 1: step-level test code generation. Specifically, the code generation system 20 receives the step to be reasoned input by the user, generates and returns the step-level script code snippet based on the step-level text description and its context.
[0143] Reasoning scenario 2: Use case level test code generation. Specifically, the code generation system 20 receives all steps and contexts of the use case to be reasoned input by the user, generates test codes for all steps in batches, and returns them.
[0144] Step 4: The code generation system 20 performs result evaluation. The detailed steps of the result evaluation are shown in Figure 12. First, the code generation system 20 creates an evaluation task in response to the task configuration operation triggered by the user. Among them, the task configuration operation includes selecting the version of the model to be tested, the task type, the evaluation data set, and the evaluation indicators. Secondly, the code generation system 20 performs verification set reasoning. Then, the code generation system 20 compares the test code generated by reasoning with the real test code in the verification set, calculates the index value of the objective evaluation index, and presents the index value of the objective evaluation index and the details of the reasoning result to the user for further subjective evaluation. Among them, the calculation formula and sample of the objective evaluation index are shown in Figure 13, which will not be repeated here.
[0145] Based on the above description, the code generation method provided by this application performs enhanced reasoning on the vector library and test method vector library by retrieving the latest local text code, thereby solving the problem that the newly added text code cannot be fed L2 Fine-Tune immediately, and solving the problem of using the newly added test method to assist in generating test code. Moreover, the method provides a set of inspection specifications for checking the training corpus, which can automatically check the corpus quantitatively, continuously iterate and manage the corpus, and significantly improve the model training effect. When training the model, the method improves the test code generation effect for one-party, two-party, and three-party libraries in specific fields through pre-training and supervised fine-tuning. The present application also provides a set of metrics for evaluating the test code generation effect, so as to achieve quantitative and objective evaluation of the effect of the test text generation code task.
[0146] Based on the aforementioned code generation method, the present application also provides a code generation system. As shown in FIG14 , the code generation system 20 includes:
[0147] Interaction module 1402, for receiving a test step description input by a user;
[0148] Retrieval module 1404, configured to retrieve a text-code pair vector library and a test method vector library based on the test step description to obtain test code snippets and test method call examples corresponding to similar test step descriptions, wherein the text-code pair vector library includes vectors formed by description text and corresponding test code, and the test method vector library includes test method vectors;
[0149] An inference module 1406 is configured to generate test code corresponding to the test step description by performing inference using a large language model (LLM) based on the test code snippets and test method call examples corresponding to the similar test step descriptions;
[0150] The interaction module 1402 is further configured to present the test code corresponding to the test step description to the user.
[0151] The interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may be modules in the aforementioned reasoning subsystem 202. For example, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may be implemented by hardware or software.
[0152] When implemented via software, the interaction module 1402, retrieval module 1404, and inference module 1406 can be applications running on a computing device, such as a computing engine. Applications can be provided as virtualization services. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, and container services. VM services can use virtualization technology to create a virtual machine (VM) resource pool on multiple physical hosts, providing users with VMs on demand. BMS services can use virtual BMS resource pools on multiple physical hosts, providing users with BMSs on demand. Container services can use virtual container resource pools on multiple physical hosts, providing users with containers on demand. A VM is a simulated virtual computer, or logically a single computer. BMS is a scalable, high-performance computing service with computing performance comparable to traditional physical machines and secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service and container service in the above-mentioned virtualization services are only specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0153] When implemented through hardware, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may include at least one computing device, such as a server. Alternatively, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0154] In some possible implementations, the system 20 further includes:
[0155] The vector library management module 1408 is used to obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library.
[0156] The vector library management module 1408 may be a module in the reasoning subsystem 202, such as a software module or a hardware module in the reasoning subsystem 202, and this embodiment does not limit this. When implemented by software, the vector library management module 1408 may be an application running on a computing device, which may be provided in the form of a virtualization service, such as a BMS service, a VM service, or a container service. When implemented by hardware, the vector library management module 1408 may include at least one computing device, such as a server. Alternatively, the vector library management module 1408 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD).
[0157] In some possible implementations, the interaction module 1402 is further configured to:
[0158] Get hints from the context of the test step description;
[0159] The reasoning module 1406 is specifically used to:
[0160] splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions;
[0161] The test code snippet spliced with the prompt and the test method call sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.
[0162] In some possible implementations, the prompt includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.
[0163] In some possible implementations, the system 20 further includes:
[0164] The training module 1410 is configured to obtain the test engineering code of the target domain, and train the universal base model according to the test engineering code of the target domain to obtain the LLM.
[0165] Training module 1410 may be a module within training subsystem 206. Training module 1410 may be implemented via software or hardware. When implemented via software, training module 1410 may be an application running on a computing device, which may be provided as a virtualized service, such as a BMS service, a VM service, or a container service. When implemented via hardware, training module 1410 may include at least one computing device, such as a server. Alternatively, training module 1410 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD).
[0166] In some possible implementations, the training module 1410 is specifically configured to:
[0167] Pre-training a universal base model based on the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model based on text-code pairs extracted from the test engineering code of the target domain to obtain the LLM; or
[0168] The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.
[0169] In some possible implementations, the system 20 further includes:
[0170] An evaluation module 1412 is configured to obtain a verification set, the verification set including text-code pairs in the target domain, the text-code pairs including descriptive text and real code snippets, use the LLM model to infer the descriptive text in the verification set to obtain generated code snippets, and determine an indicator value of an evaluation indicator based on the generated code snippets and the real code snippets, the evaluation indicator including at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter value recall rate, and test parameter value precision rate;
[0171] The interaction module 1402 is further configured to present the indicator value of the evaluation indicator to the user.
[0172] Evaluation module 1412 may be a module within training subsystem 206. Evaluation module 1412 may be implemented via software or hardware. When implemented via software, evaluation module 1412 may be an application running on a computing device, which may be provided as a virtualized service, such as a BMS service, a VM service, or a container service. When implemented via hardware, evaluation module 1412 may include at least one computing device, such as a server. Alternatively, evaluation module 1412 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD).
[0173] In some possible implementations, the interaction module 1402 is further configured to:
[0174] The generated code snippet and the real code snippet are presented to the user.
[0175] This application also provides a computing device 1500. As shown in Figure 15, computing device 1500 includes a bus 1502, a processor 1504, a memory 1506, and a communication interface 1508. Processor 1504, memory 1506, and communication interface 1508 communicate with each other via bus 1502. Computing device 1500 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1500.
[0176] Bus 1502 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG15 shows a single bus line, but this does not imply a single bus or type of bus. Bus 1502 may include a path for transmitting information between various components of computing device 1500 (e.g., memory 1506, processor 1504, and communication interface 1508).
[0177] The processor 1504 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0178] The memory 1506 may include a volatile memory, such as a random access memory (RAM). The memory 1506 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory 1506 stores an executable program code, and the processor 1504 executes the executable program code to implement the aforementioned code generation method. Specifically, the memory 1506 stores instructions for the code generation system 20 to execute the code generation method.
[0179] The communication interface 1508 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1500 and other devices or a communication network.
[0180] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0181] As shown in Figure 16, the computing device cluster includes at least one computing device 1500. The memory 1506 of one or more computing devices 1500 in the computing device cluster may store the same code generation system 20 for executing instructions of the code generation method.
[0182] In some possible implementations, one or more computing devices 1500 in the computing device cluster may also be used to execute some of the instructions of the code generation system 20 for executing the code generation method. In other words, the combination of one or more computing devices 1500 may jointly execute the instructions of the code generation system 20 for executing the code generation method.
[0183] It should be noted that the memories 1506 in different computing devices 1500 in the computing device cluster may store different instructions for executing partial functions of the code generation system.
[0184] Figure 17 shows a possible implementation. As shown in Figure 17, two computing devices 1500A and 1500B are connected via a communication interface 1508. The memory in computing device 1500A stores instructions for executing the functions of interaction module 1402 and retrieval module 1404. The memory in computing device 1500B stores instructions for executing the functions of reasoning module 1406. In other words, the memory 1506 of computing devices 1500A and 1500B jointly stores instructions for the code generation system 20 to execute the code generation method. Further, computing device 1500A, computing device 1500B or other computing devices can also store instructions for executing the functions of vector library management module 1408, training module 1410, and evaluation module 1412.
[0185] The connection mode between the computing device clusters shown in FIG17 may be considered to be due to the fact that the code generation method provided in this application requires a lot of computing power for reasoning. Therefore, it is considered to delegate the functions implemented by the reasoning module 1406 to the computing device 1500B.
[0186] It should be understood that the functionality of the computing device 1500A shown in FIG17 may also be implemented by multiple computing devices 1500. Similarly, the functionality of the computing device 1500B may also be implemented by multiple computing devices 1500.
[0187] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), among others. FIG. 18 illustrates a possible implementation. As shown in FIG. 18 , two computing devices 1500C and 1500D are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1506 in the computing device 1500C stores instructions for executing the functions of the interaction module 1402 and the retrieval module 1404. Simultaneously, the memory 1506 in the computing device 1500D stores instructions for executing the functions of the reasoning module 1406. Furthermore, the computing device 1500C, the computing device 1500D, or other computing devices may also store instructions for executing the functions of the vector library management module 1408, the training module 1410, and the evaluation module 1412.
[0188] The connection method between the computing device clusters shown in Figure 18 can be that considering that the code generation method provided by this application requires a lot of computing power for reasoning, it is considered to hand over the functions implemented by the reasoning module 1406 to the computing device 1500D for execution.
[0189] It should be understood that the functionality of the computing device 1500C shown in FIG18 may also be accomplished by multiple computing devices 1500. Similarly, the functionality of the computing device 1500D may also be accomplished by multiple computing devices 1500.
[0190] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned code generation system 20 for executing the code generation method.
[0191] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the aforementioned code generation method.
[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A code generation method, characterized in that: The method is performed by a code generation system, wherein the code generation system is used to automatically generate test code, and the method includes: Receive test step descriptions from user input; According to the test step description, a text code pair vector library and a test method vector library are retrieved to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes a description text and a vector formed by the corresponding test code, and the test method vector library includes a test method vector; According to the test code snippets and test method call examples corresponding to the similar test step descriptions, reasoning is performed through the large language model LLM to generate the test code corresponding to the test step description; The test code corresponding to the test step description is presented to the user.
2. The method according to claim 1, characterized in that The method further comprises: Obtain incremental test text code pairs and incremental test methods from the target domain's code repository or integrated development and testing environment; The incremental test text code pairs are vectorized and added into the text code pair vector library, and the incremental test methods are vectorized and added into the test method vector library.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Taking hints from the context of the test step description; The method of performing reasoning based on the test code snippets and test method call examples corresponding to the similar test step descriptions through a large language model (LLM) to obtain the test code corresponding to the test step descriptions includes: splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions; The test code snippet spliced with the prompt and the test method calling sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.
4. The method according to claim 3, characterized in that The hint includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Obtain the test engineering code of the target domain; The LLM is obtained by training a general base model according to the test engineering code of the target domain.
6. The method according to claim 5, characterized in that The step of training a base model according to the test engineering code of the target domain to obtain the LLM comprises: Pre-training a universal base model according to the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model according to text code pairs extracted from the test engineering code of the target domain to obtain the LLM; or The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.
7. The method according to claim 5 or 6, characterized in that: The method further comprises: Acquire a verification set, wherein the verification set includes text-code pairs of the target domain, and the text-code pairs include description text and real code snippets; Use the LLM model to infer the description text in the verification set to obtain a generated code snippet; Determine an index value of an evaluation index according to the generated code snippet and the real code snippet, wherein the evaluation index includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate, and test parameter assignment precision rate; The indicator value of the evaluation indicator is presented to the user.
8. The method according to claim 7, characterized in that The method further comprises: The generated code snippet and the actual code snippet are presented to the user.
9. The method according to any one of claims 1 to 8, characterized in that: The test step description includes a description based on a natural language or a description based on a domain specific language.
10. A code generation system, characterized in that: The code generation system is used to automatically generate test code, and the system includes: An interactive module, used for receiving a test step description input by a user; A retrieval module is used to retrieve a text code pair vector library and a test method vector library according to the test step description to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes a description text and a vector formed by the corresponding test code, and the test method vector library includes a test method vector; An inference module, configured to generate test code corresponding to the test step description by performing inference through a large language model (LLM) based on the test code snippets and test method call examples corresponding to the similar test step descriptions; The interaction module is further used to present the test code corresponding to the test step description to the user.
11. The system according to claim 10, characterized in that The system further comprises: The vector library management module is used to obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library.
12. The system according to claim 10 or 11, characterized in that The interaction module is also used for: Taking hints from the context of the test step description; The reasoning module is specifically used for: splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions; The test code snippet spliced with the prompt and the test method calling sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.
13. The system according to claim 12, characterized in that The hint includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.
14. The system according to any one of claims 10 to 13, characterized in that: The system further comprises: The training module is used to obtain the test engineering code of the target field, train the universal base model according to the test engineering code of the target field, and obtain the LLM.
15. The system according to claim 14, characterized in that The training module is specifically used for: Pre-training a universal base model according to the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model according to text code pairs extracted from the test engineering code of the target domain to obtain the LLM; or, The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.
16. The system according to claim 14 or 15, characterized in that The system further comprises: An evaluation module is used to obtain a verification set, wherein the verification set includes text-code pairs of the target domain, wherein the text-code pairs include description texts and real code snippets, and use the LLM model to infer the description texts in the verification set to obtain generated code snippets, and determine the index values of evaluation indicators according to the generated code snippets and the real code snippets, wherein the evaluation indicators include at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter value recall rate, and test parameter value precision rate; The interaction module is further used to present the indicator value of the evaluation indicator to the user.
17. The system according to claim 16, characterized in that The interaction module is also used for: The generated code snippet and the actual code snippet are presented to the user.
18. The system according to any one of claims 10 to 17, characterized in that The test step description includes a description based on a natural language or a description based on a domain specific language.
19. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the computing device cluster executes the code generation method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that: The method comprises computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.
21. A computer program product, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Test case processing method and server
CN107704392A
Information processing method and device and storage medium
CN116483372A
Code generation method and device
CN116719520A
Test case generation method and device, electronic equipment and storage medium
CN116955210A
Test case generation method and device, terminal equipment and storage medium
CN116991711A
Cited By
Method and system for intelligent computing center cloud platform to realize agent self-evolution through computing power
CN120276712A
Code generation method and system based on large model
CN120763072A
Platform and method for automatically testing and generating kernel export function of operating system
CN120803962A
Large code model training method and electronic equipment
CN121187569A