Code generation method and related equipment

By combining search and large language model (LLM) inference, the problem of low efficiency in testing code automation development in the existing technology in integrated testing and system testing is solved, and efficient automation development of complex test scenarios is achieved.

CN120029897APending Publication Date: 2025-05-23HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311602430.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2023-11-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

现有技术在集成测试和系统测试中测试代码的自动化开发效率较低,难以满足复杂测试场景的业务需求。

Method used

By combining search and large language model (LLM) inference, search text code to vector library and test method vector library to obtain test code snippets and test method call samples corresponding to similar test step descriptions, and generate test code through LLM inference.

Benefits of technology

It realizes automatic generation of complex test code, improving the efficiency of automated development of complex test scenarios (such as integration testing, system testing).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029897A_ABST
    Figure CN120029897A_ABST
Patent Text Reader

Abstract

The invention provides a code generation method which is executed by a code generation system, the code generation system is used for automatically generating a test code, and the method comprises the following steps: receiving a test step description input by a user, retrieving a text code pair vector library and a test method vector library according to the test step description, the text code pair vector library comprises description texts and vectors formed by corresponding test codes, and the test method vector library comprises vectors formed by the test method calling examples; and according to the test code snippets and the test methods corresponding to the similar test step descriptions, reasoning through a large language model (LLM), generating test codes corresponding to the test step descriptions, and presenting the test codes corresponding to the test step descriptions to a user. According to the method, retrieval and LLM reasoning are combined, the complex test code can be automatically generated, and the automatic development efficiency of the complex test scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 17, 2023, with application number 202311544947.4 and invention name “A code generation method and related equipment”, the entire contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a code generation method, system, computing device cluster, computer-readable storage medium, and computer program product. Background Art

[0003] From software development to release, it is usually necessary to undergo a series of tests, including but not limited to unit testing (UT), integration testing, and system testing. Unit testing, integration testing, or system testing can be implemented by executing corresponding test codes (such as test script codes).

[0004] For unit testing, the automated development technology of test code is relatively mature. With the rise of the general Large Language Model (LLM), it is possible to use the general LLM to generate function-level unit test code. However, since integration testing or system testing is more complex, the automated development efficiency of test code is low and it is difficult to meet business needs. Summary of the invention

[0005] The present application provides a code generation method, which combines retrieval and LLM reasoning. The method first retrieves the text code pair vector library and the test method vector library to obtain the test code snippets and test method call samples corresponding to the similar test step descriptions. Based on the test code snippets and test method call samples, LLM reasoning is performed instead of reasoning at the function granularity, so that complex test codes can be automatically generated and the efficiency of automated development of complex test scenarios can be improved. The present application also provides a code generation system, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above method.

[0006] In the first aspect, the present application provides a code generation method. The method can be executed by a code generation system (or referred to as a code generation platform, a code generation tool). The code generation system is used to automatically generate test code. Among them, the code generation system can be a software system, and the software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. It should be noted that the software system can be an independent software package or provide an interface for users to use in the form of a cloud service, and the software system can also be integrated in other software for users to use in the form of a plug-in, for example, the software system can be a plug-in for an integrated development environment (IDE). In some examples, the code generation system can also be a hardware system, such as a computing device cluster with code generation capability, and the computing device cluster executes the code generation method of the present application when it is running.

[0007] Specifically, the code generation system receives a test step description input by a user, and then retrieves a text code pair vector library and a test method vector library according to the test step description to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes vectors formed by description texts and corresponding test codes, and the test method vector library includes test method vectors. Then, the code generation system generates test code corresponding to the test step description by performing reasoning through a large language model (LLM) based on the test code snippets and test method call samples corresponding to similar test step descriptions. The code generation system presents the test code corresponding to the test step description to the user.

[0008] This method combines retrieval and LLM reasoning. It first searches the text code pair vector library and the test method vector library to obtain test code snippets and test method call samples corresponding to similar test step descriptions (for example, test method call samples with similar semantics). Based on the test code snippets and test method call samples, LLM reasoning is performed instead of reasoning at the function granularity. This can automatically generate complex test code (such as system test code) and improve the efficiency of automated development of complex test scenarios (such as integration testing and system testing).

[0009] In some possible implementations, the code generation system can obtain incremental test text code pairs and incremental test methods from a code repository or an integrated development and testing environment in the target domain. The test method includes at least one of a method definition, a method signature, a parameter description, a method body, and a call example. The code generation system can then vectorize the incremental test text code pairs and add them to a text code pair vector library, and vectorize the incremental test methods and add them to a test method vector library.

[0010] By adopting the incremental update mechanism, the text code pair vector library can store the latest test step descriptions and test text code pairs formed by the test code, and the test method vector library can store the latest test methods. This can ensure the comprehensiveness, accuracy and timeliness of the retrieval results, and the test code can be generated based on the retrieval results, which can achieve better results and ensure the quality of the generated test code.

[0011] In some possible implementations, the code generation system can also obtain hints from the context of the test step description. Accordingly, the code generation system can splice the hints with the test code snippets and test method call samples corresponding to similar test step descriptions, and then input the spliced ​​hinted test code snippets and test methods into the LLM for reasoning to obtain the test code corresponding to the test step description.

[0012] This method can improve the accuracy of reasoning by combining hints in the context, thereby improving the quality of the test code generated by reasoning.

[0013] In some possible implementations, the prompt includes at least one of global information, use case level information, step level information, or cross-use case context information in the context. Among them, global information may include product information, file information, and test language framework information. Product information includes product lines and product R&D units, file information includes file names, Git repositories, and file path information, use case level information includes at least one of test case attribute information, import libraries, use case base class information, and variable initialization definitions, use case attribute information may include at least one of use case class name, use case name, test type, test activity, characteristics, test environment type, preconditions, and test step descriptions, import libraries include at least one of import class signatures, comments or import method signatures, and comments, and use case base class information includes at least one of base class signatures, comments or base class method signatures, and comments. Step-level information includes at least one of step text descriptions, step test script code blocks, and preceding step codes. Cross-use case context information includes at least one of test suite information, environment information, and configuration information. Test suite information includes at least one of test suite attribute definitions, test suite method signatures, and comments.

[0014] By splicing at least one of the above global information, use case level information, step level information, and cross-use case context information, richer information can be provided to LLM, thereby improving the accuracy of LLM reasoning in generating test code.

[0015] In some possible implementations, the code generation system can also obtain the test engineering code of the target domain, and then train the general base model according to the test engineering code of the target domain to obtain the LLM. In this way, the trained LLM can be used to automatically generate test code for a specific domain (such as the target domain), rather than being limited to generating test code for general Internet scenarios, which has high usability.

[0016] In some possible implementations, when the code generation system trains the base model to obtain the LLM, the general base model may be pre-trained according to the test engineering code of the target domain to obtain the test code generation pre-trained model, and the test code generation pre-trained model may be fine-tuned in a supervised manner according to the text code pairs extracted from the test engineering code of the target domain to obtain the LLM. Alternatively, the code generation system may fine-tune the general base model in a supervised manner according to the text code pairs extracted from the test engineering code of the target domain to obtain the LLM.

[0017] This method can improve the model training efficiency by pre-training a common base model and then performing supervised fine-tuning on the pre-trained model, or directly performing supervised fine-tuning on the common base model.

[0018] In some possible implementations, the code generation system can also obtain a verification set, which includes text code pairs in the target domain, and the text code pairs include description texts and real code snippets. Then the code generation system can use the LLM model to reason about the description texts in the verification set to obtain generated code snippets. Accordingly, the code generation system can determine the index value of the evaluation index based on the generated code snippets and the real code snippets. Among them, the evaluation index includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate or test parameter assignment precision rate. The code generation system can present the index value of the evaluation index to the user.

[0019] This method also supports comparing the inference results with the ground truth (GT) and calculating the index values ​​of the evaluation indicators, thereby achieving objective evaluation of LLM and providing a reference for users to decide whether to accept the test code generated by LLM reasoning.

[0020] In some possible implementations, the code generation system may also present the generated code snippet and the real code snippet to the user, so as to support the user to directly compare the test code generated by reasoning with the real test code, thereby realizing the subjective evaluation of LLM and providing a reference for the user to decide whether to accept the test code generated by LLM reasoning.

[0021] In some possible implementations, the test step description includes a description based on a natural language or a description based on a domain specific language, wherein the domain specific language, also known as a domain-specific language, refers to a computer language that focuses on a certain application domain.

[0022] This method supports the generation of test code based on natural language or domain-specific language, without the need to solidify the test step description, and has high usability and flexibility.

[0023] In some possible implementations, the code generation system receives a step-level test step description. Accordingly, the code generation system can retrieve a text code pair vector library and a test method vector library according to the step-level test step description, obtain a test code snippet and a test method call sample corresponding to a similar step-level test step description, and generate a step-level test code through LLM reasoning based on the test code snippet and the test method call sample. In this way, the test code can be generated step by step, and manual intervention or adjustment can be made during the test code generation process to ensure the accuracy of the generated test code.

[0024] In some possible implementations, the code generation system receives a use case level test step description, which may be a complete test step description. Accordingly, the code generation system may retrieve a text code pair vector library and a test method vector library based on the use case level test step description, obtain test code snippets and test method call examples corresponding to similar step level test step descriptions, and generate use case level test code through LLM reasoning based on the test code snippets and test method call examples.

[0025] In this way, the test codes for all steps can be generated in batches and returned to the user, thereby improving the efficiency of test code generation.

[0026] In some possible implementations, the code generation system can also perform a specification check on the training corpus. Examples of valid specification check items include the following categories: uniqueness check, validity check, integrity check, consistency check, and security check. Each category can also include several check items. Taking consistency check as an example, consistency check can be subdivided into text code consistency (also known as text script consistency check), use case text quality check, test method quality check, test method density check, and test method style check. Furthermore, text code consistency check can be further subdivided into the following check items: text similarity, inconsistent test code implementation; similar test code implementation, inconsistent text description.

[0027] This method can improve the quality of training data by performing the above-mentioned standard check and other data quality engineering on the training corpus. Training LLM based on high-quality training data can improve training efficiency and training effect.

[0028] In a second aspect, the present application provides a code generation system. The code generation system is used to automatically generate test code, and the system includes:

[0029] An interactive module, used for receiving a test step description input by a user;

[0030] A retrieval module is used to retrieve a text code pair vector library and a test method vector library according to the test step description to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes a description text and a vector formed by the corresponding test code, and the test method vector library includes a test method vector;

[0031] An inference module, configured to generate test code corresponding to the test step description by performing inference through a large language model (LLM) based on the test code snippets and test method call examples corresponding to the similar test step descriptions;

[0032] The interaction module is further used to present the test code corresponding to the test step description to the user.

[0033] In some possible implementations, the system further includes:

[0034] The vector library management module is used to obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library.

[0035] In some possible implementations, the interaction module is further configured to:

[0036] Taking hints from the context of the test step description;

[0037] The reasoning module is specifically used for:

[0038] splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions;

[0039] The test code snippet spliced ​​with the prompt and the test method calling sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.

[0040] In some possible implementations, the prompt includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.

[0041] In some possible implementations, the system further includes:

[0042] The training module is used to obtain the test engineering code of the target field, train the universal base model according to the test engineering code of the target field, and obtain the LLM.

[0043] In some possible implementations, the training module is specifically used to:

[0044] Pre-training a universal base model according to the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model according to text code pairs extracted from the test engineering code of the target domain to obtain the LLM; or

[0045] The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.

[0046] In some possible implementations, the system further includes:

[0047] An evaluation module is used to obtain a verification set, wherein the verification set includes text-code pairs of the target domain, wherein the text-code pairs include description texts and real code snippets, and use the LLM model to infer the description texts in the verification set to obtain generated code snippets, and determine the index values ​​of evaluation indicators according to the generated code snippets and the real code snippets, wherein the evaluation indicators include at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter value recall rate, and test parameter value precision rate;

[0048] The interaction module is further used to present the indicator value of the evaluation indicator to the user.

[0049] In some possible implementations, the interaction module is further configured to:

[0050] The generated code snippet and the actual code snippet are presented to the user.

[0051] In some possible implementations, the test step description includes a description based on a natural language or a description based on a domain specific language.

[0052] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory, so that the computing device or the computing device cluster executes the code generation method described in the first aspect or any implementation of the first aspect.

[0053] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, wherein the instructions instruct a computing device or a computing device cluster to execute the code generation method described in the above-mentioned first aspect or any one of the implementations of the first aspect.

[0054] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the code generation method described in the first aspect or any one of the implementations of the first aspect.

[0055] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical method of the embodiments of the present application, the drawings required for use in the embodiments are briefly introduced below.

[0057] Figure 1 A schematic diagram of a test framework for a system test provided in this application;

[0058] Figure 2 A schematic diagram of the architecture of a code generation system provided for this application;

[0059] Figure 3 A schematic diagram of a training and reasoning process provided for this application;

[0060] Figure 4 A schematic diagram of a process for enhancing reasoning of a vector library and a test method vector library by a retrieval text code provided in this application;

[0061] Figure 5 A schematic diagram of an evaluation indicator provided for this application;

[0062] Figure 6 A flowchart of a code generation method provided for this application;

[0063] Figure 7 A schematic diagram of obtaining a prompt through a prompting project provided by the present application;

[0064] Figure 8 A schematic diagram of a test code corpus checking specification provided for this application;

[0065] Fig. 9 A flowchart of a data inspection and cleaning tool provided for this application;

[0066] Fig.10A schematic diagram of a data inspection and cleaning result provided for this application;

[0067] Fig.11 A schematic diagram of a reasoning tool provided for this application;

[0068] Fig.12 A schematic diagram of a result evaluation process provided for this application;

[0069] Fig.13 A calculation formula and sample diagram of an objective evaluation indicator provided for this application;

[0070] Fig.14 A schematic diagram of the structure of a code generation system provided for this application;

[0071] Fig.15 A schematic diagram of the structure of a computing device provided for this application;

[0072] Fig.16 A schematic diagram of the structure of a computing device cluster provided for this application;

[0073] Fig.17 A schematic diagram of the structure of another computing device cluster provided for this application;

[0074] Fig.18 A schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION

[0075] The terms "first" and "second" in the embodiments of the present application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.

[0076] First, some technical terms involved in the embodiments of the present application are introduced.

[0077] Unit testing, also known as module testing, is a test that verifies the correctness of the smallest unit of software design (such as a "program module" or "program unit"). The purpose of unit testing is to check whether each program unit can correctly implement the requirements of the functions, performance, interfaces, and design constraints in the detailed design specifications, and to discover various errors that may exist in each unit. Unit testing requires designing test cases based on the internal structure of the program. Multiple program units can be independently unit tested in parallel.

[0078] Integration testing, also known as assembly testing, is usually based on unit testing. It tests program units (such as software units, components, and subsystems) in an orderly and incremental manner to check whether there are any problems with the interfaces between program units such as software units, components, and subsystems. Integration testing is to check the interface relationship of program units or components, and gradually integrate them into program components or software that meet the requirements of the outline design.

[0079] System testing is to combine the software that has passed the integration test as a part of the system with system elements such as hardware, supporting software, data and platform, and test the quality attributes of the system under test, such as system functionality, performance, reliability, security, resilience, maintainability, etc. in a simulated or real environment, so as to discover potential problems and defects in the software, detect code logic and operating results, and verify whether the system meets user needs.

[0080] Compared with unit testing, the testing process of integration testing and system testing is more complicated. The following is an example of system testing.

[0081] System testing needs to be run on a real hardware environment or simulated hardware environment (for ease of description, it can also be referred to as a system testing environment) that integrates many software under test and platform components. Therefore, the test framework of system testing is usually much more complex than that of unit testing. Figure 1 The schematic diagram of a test framework for a system test shown in the figure includes a driver layer, a support layer, a business layer and an application layer, wherein the driver layer is also called the underlying communication and message parsing layer, which is used to realize communication interaction with the device, command result parsing, and automated script case execution framework, and the support layer is also called the product support layer, which usually includes the class library of the tested product (such as the test object basic library obtained after analyzing, designing, and coding each tested system in the product using the test object model) and the reusable public basic library of some products. The business layer can include the product command line encapsulation layer (command action word, CAW), the business model layer (business action word, BAW), and the test model layer. The application layer includes the test case description layer. The test method is encapsulated layer by layer, and finally forms the test code of the application layer (usually in the form of a script, also called a test script code), which can be run on the system test environment, and the test steps and checkpoints are instantiated by calling BAW / CAW.

[0082] The complexity of the above test framework leads to low efficiency in system test automation code development, which is difficult to meet the automation requirements of new use cases. Moreover, the test needs to maintain the existing functions without damage. As the functions increase, the workload of continuous testing continues to increase, and automation is needed to solve the coverage test of existing functions. With the increasing workload of business testing, it is difficult to guarantee the manpower that the business team can invest in automation, and automation problems cannot be solved by increasing manpower alone.

[0083] Test code can be automatically generated based on general LLM, but it is limited to function-level unit test scenarios. In complex test scenarios such as system testing, the code covered by a test case is X hundreds to X thousands or even X tens of thousands of functions, and system test code cannot be generated through a full white box context.

[0084] In view of this, the present application provides a code generation method. The method can be executed by a code generation system (or referred to as a code generation platform, a code generation tool). The code generation system is used to automatically generate test code. Among them, the code generation system can be a software system, and the software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. It should be noted that the software system can be an independent software package or provide an interface for users to use in the form of a cloud service, and the software system can also be integrated in other software for users to use in the form of a plug-in, for example, the software system can be a plug-in for an integrated development environment (IDE). In some examples, the code generation system can also be a hardware system, such as a computing device cluster with code generation capabilities, and the computing device cluster executes the code generation method of the present application when it is running.

[0085] Specifically, the code generation system receives a test step description input by a user, and then retrieves a text code pair vector library and a test method vector library according to the test step description to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes vectors formed by description texts and corresponding test codes, and the test method vector library includes test method vectors. Then, the code generation system generates test code corresponding to the test step description by performing reasoning through a large language model (LLM) based on the test code snippets and test method call samples corresponding to similar test step descriptions. The code generation system presents the test code corresponding to the test step description to the user.

[0086] This method combines retrieval and LLM reasoning. It first searches the text code pair vector library and the test method vector library to obtain test code snippets and test method call samples corresponding to similar test step descriptions (for example, test method call samples with similar semantics). Based on the test code snippets and test method call samples, LLM reasoning is performed instead of reasoning at the function granularity. This can automatically generate complex test code (such as system test code) and improve the efficiency of automated development of complex test scenarios (such as integration testing and system testing).

[0087] Moreover, the method can generate test methods for local first-party, second-party, and third-party libraries and test codes for custom test frameworks for business contexts in specific domains through vector libraries in specific domains, such as text code pair vector libraries in the target domain and test method vector libraries in the target domain. The method is not limited to generating test codes for general Internet scenarios and general test frameworks and has high availability.

[0088] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.

[0089] See also Figure 2 The schematic diagram of the architecture of a code generation system shown in the figure, the code generation system 20 includes a reasoning subsystem 202, and the reasoning subsystem 202 generates test code through LLM reasoning. Further, the code generation system 20 may also include a data processing subsystem 204, a training subsystem 206, and an evaluation subsystem 208. Among them, the data processing subsystem 204 is used to implement data quality engineering (or data engineering), and the training subsystem 206 is used to perform model training based on the data processed by the data processing subsystem 204 to obtain LLM. The reasoning subsystem 202 is used to perform reasoning based on LLM and generate test code. The evaluation subsystem 208 is used to perform reasoning through LLM and perform evaluation according to the reasoning results.

[0090] The overall processing flow is described in detail below in conjunction with the system architecture of the code generation system 20.

[0091] In the first stage, the data processing subsystem 204 performs data quality engineering. Specifically, the data processing subsystem 204 obtains training corpus, processes the training corpus, and obtains training data. Among them, the data processing subsystem 204 can be a unified corpus quality inspection and cleaning tool for different data sources, and the data processing subsystem 204 can define at least one of the training data format standards or inspection specifications. In this embodiment, the training data format standard can include at least one of the test code raw-code corpus data format standard, the test text code pair corpus data format standard, or the test method signature annotation-test code corpus data format standard. Training sets of different training data format standards can be used to train different LLMs. For example, the training set of the test code raw-code corpus data format standard is used to train L1 LLM, and L1LLM can be a LLM obtained by training (such as pre-training) a general LLM (general base model, also called L0 LLM) for generating test code, also called a test code generation pre-training model, based on which, the above-mentioned test code raw-code corpus data format standard can also be called L1 test code raw-code corpus data format standard. For another example, the test text code pair corpus data format standard or the test method signature annotation-test code corpus data format standard is used to train the L2 LLM, and the L2 LLM can be an LLM directly trained from a general LLM for generating test codes, or an LLM trained from the above-mentioned L1 LLM for generating test codes. Among them, supervised fine-tuning (SFT) can be used to train the L2 LLM. Based on this, the above-mentioned test text code pair corpus data format standard and the test method signature annotation-test code corpus data format standard can also be referred to as the L2 test text code pair corpus data format standard and the L2 test method signature annotation-test code corpus data format standard. The inspection specifications may include corpus quality inspection and cleaning specifications, such as L1 / L2 corpus quality inspection and cleaning specifications.

[0092] The data processing subsystem 204 can extract training data from the original test engineering code based on the training data format standard. Furthermore, the data processing subsystem 204 can check the training data based on the inspection specification. For the data that meets the inspection specification, the data processing subsystem 204 can also clean, optimize, and evaluate the data quality characteristics.

[0093] In some possible implementations, the data processing subsystem 204 may also reserve data as a validation set for subsequent evaluation. Based on this, the validation set may also be referred to as an evaluation set.

[0094] In the second stage, the training subsystem 206 performs model training. The training subsystem 206 can obtain test engineering codes in a specific field, such as test engineering codes in the target field, train the general base model according to the test engineering codes in the target field, and obtain an LLM (also called a domain test code LLM, which can be the above-mentioned L2 LLM) that can automatically generate test codes in the target field.

[0095] like Figure 3 As shown, the training subsystem 206 selects an L0 LLM (universal base model) suitable for completing the task of text-to-code conversion (e.g., natural language text-to-code conversion), obtains one or more test engineering codes in a specific field, and trains the L0 LLM in a pre-training manner (e.g., incremental pre-training) according to the test engineering code (e.g., cleaned and optimized test engineering code) to obtain one or more pre-trained models in a specific field, which are also called test code generation pre-training models and belong to L1 LLM. On this basis, the training subsystem 206 extracts text code pairs from the test engineering codes of the corresponding one or more specific fields, which can be test case text-test code snippet pairs, and performs supervised fine-tuning on the L1 LLM to obtain an LLM that can automatically generate test code. The LLM can generate test code according to the test text (e.g., test step description) in a specific field, which belongs to L2 LLM. Optionally, the training subsystem 206 can also directly perform SFT on the basis of the L0 LLM to obtain an L2 LLM for subsequent reasoning. Taking the training effect into consideration, the training subsystem 206 may also perform prompt engineering. Accordingly, when performing model training, it may combine the prompts obtained from the prompt engineering, such as at least one of the use case global info, dependency library info, and step info prompts, to perform supervised fine-tuning.

[0096] In the third stage, the reasoning subsystem 202 performs enhanced reasoning generation based on the retrieval text code pair vector library and the test method vector library, also known as retrieve augmented generation (RAG). The reasoning subsystem 202 may include a front-end plug-in, a back-end intelligent agent (referred to as an agent), a vector library, and an LLM reasoning service cluster. The vector library includes the aforementioned text code pair vector library and the test method vector library.

[0097] like Figure 4As shown, in the online reasoning stage, the user can input the test step description to be reasoned (in some cases, it can also include the test case intention) through the interface of the front-end plug-in or the agent. Accordingly, the agent retrieves the text code pair vector library and the test method vector library to obtain the test code snippet corresponding to the similar test step description (for example, the test code snippet corresponding to the test step description with the top K similarity ranking) and the test method call sample (use case), wherein the test method call sample can be a call sample with similar semantics (for example, a call sample with the top K semantic similarity ranking). Then the agent can also combine other contexts of the test case, instantiate the prompt of the prompt engineering template, call the LLM reasoning service interface (which can be L0 LLM, or L1 LLM or L2 LLM), and obtain the test code generated by LLM reasoning. Furthermore, the agent can also perform post-processing such as truncation, de-repeating, and indentation problem repair on the generated test code. The agent returns the test code to the user through the interface of the front-end plug-in or the agent. The user can give feedback on the generated test code, such as accepting the test code, modifying the test code, or rejecting the test code. The reasoning subsystem 202 may store the user's feedback on the generated test code in a feedback database (feedback DB).

[0098] Among them, the reasoning subsystem 202 sends a query request to the retrieval retrieval module for similarity retrieval. After obtaining the retrieval results, it can also determine whether the similarity between the retrieval results and the test step description (such as the text to be inferred, including but not limited to the test step description to be inferred) meets the requirements, and decide the subsequent processing according to the judgment result. For example, if the similarity is higher than the set value, the reasoning subsystem 202 can splice the retrieval prompt and perform reasoning through LLM. For another example, if the similarity is lower than the set value, the reasoning subsystem 202 can discard the retrieval result.

[0099] In some possible implementations, the reasoning subsystem 202 can also be connected to a code repository, such as a domain code repository, a personal test project code repository, or an integrated development and testing environment. Accordingly, in the offline stage, the reasoning subsystem 202 can also obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library. It should be noted that the above-mentioned incremental data can also be added to the SFT DB for performing SFT on the LLM. In addition, the reasoning subsystem 202 also supports data deletion or aging, including automatic data aging, or deleting data in response to a deletion operation actively triggered by the user.

[0100] It should be noted that the reasoning subsystem 202 can perform reasoning on the verification set before the LLM goes online to evaluate the LLM. This reasoning process is also called model evaluation reasoning. The reasoning subsystem 202 can also perform reasoning after the LLM goes online to evaluate the running experience of the LLM. For example, the reasoning subsystem 202 can generate test code by reasoning, and then combine the generated code and the real code to count the code retention results.

[0101] In the fourth stage, the evaluation subsystem 208 performs reasoning through LLM and performs evaluation based on the reasoning results. In some possible implementations, the evaluation subsystem 208 can obtain a verification set, which can include text-code pairs in a specific field (such as a target field), and the evaluation subsystem 208 can perform objective evaluation or subjective evaluation on the reasoning results of the description text in the verification set (code snippets generated based on the description text, referred to as generated code snippets) through the reasoning subsystem 202.

[0102] The objective evaluation process includes: the evaluation subsystem 208 determines the index value of the evaluation index according to the generated code snippet and the real code snippet in the verification set. The evaluation index includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate, and test parameter assignment precision rate. In some possible implementations, the index value of the evaluation index can be obtained by Figure 5 The evaluation subsystem 208 can display the index values ​​of the above evaluation indexes to the user.

[0103] The subjective evaluation process includes: the evaluation subsystem 208 presents the generated code snippet and the real code snippet to the user, so that the user can compare the generated code snippet with the real code snippet to perform subjective evaluation.

[0104] Based on the aforementioned code generation system 20, the present application provides a code generation method. The code generation method of the present application is described in detail below in conjunction with an embodiment.

[0105] See also Figure 6 A flowchart of a code generation method is shown, the method comprising the following steps:

[0106] S602: The code generation system 20 receives a test step description input by a user.

[0107] The test step description is information describing the test steps. In the present embodiment, the test step description may be a description based on a natural language (NL) or a description based on a domain-specific language (DSL). Among them, a domain-specific language, also known as a domain-specific language, refers to a computer language that focuses on a certain application domain. Different from ordinary cross-domain general-purpose programming languages ​​(GPL), domain-specific languages ​​are usually used in certain specific fields, such as Hyper Text Markup Language (HTML) used to display web pages.

[0108] The code generation system 20 supports at least one input method to input the test step description. For example, the code generation system 20 can present a graphical interface or a command line interface to the user, and the user inputs the test step description in the form of text input through the graphical interface or the command line interface, and accordingly, the code generation system 20 can receive the test step description in text form. For another example, the code generation system 20 can provide a voice input control and a voice input switching instruction. When the voice input control or the voice input switching instruction is triggered, the user can input the test step description in the form of voice input, and accordingly, the code generation system 20 can receive the test step description in voice form. Furthermore, the code generation system 20 can also convert the test step description in voice form into the test step description in text form through voice recognition or voice-to-text technology.

[0109] In the scheme of generating test code through LLM reasoning, this embodiment provides multiple scenarios. One scenario is that the user inputs a complete test step description, and the code generation system 20 receives the complete test step description input by the user, so as to generate the test code at one time. Another scenario is that the user inputs a partial test step description, and the code generation system 20 receives the partial test step description input by the user, so as to generate the test code step by step. For example, the user can first input the test step description of the first step, and when the code generation system 20 returns the corresponding test code, the user can then input the test step description of the second step, and continue to generate the test code of the second step. And so on, which will not be repeated here.

[0110] When inputting the test step description, the user may input the test step description in the form of comments in the code file. Alternatively, the user may input the test step description in the form of an independent file. This embodiment does not limit this.

[0111] S604. The code generation system 20 retrieves the text code pair vector library and the test method vector library according to the test step description, and obtains the test code snippets and test method calling samples corresponding to similar test step descriptions.

[0112] The text code pair vector library includes a vector formed by a description text and a corresponding test code, and the test method vector library includes a test method vector. Wherein, the description text refers to the text of the test step description, which can also be called a test text, including but not limited to a test case text. The test method may include at least one of a method definition, a method signature, a parameter description, a method body, and a call example. The above vectors in the vector library can be used as an index, for example, the vectors in the text code pair vector library are used to index the corresponding text code pairs, and the vectors in the test method vector library are used to index the corresponding test methods. In order to realize the automatic generation of test codes for specific fields, the text code pair vector library can be a text code pair vector library for specific fields, and the test method vector library can be a test method vector library for specific fields.

[0113] Among them, the text code pair vector library and the test method vector library can use the incremental update mechanism to update the database. Specifically, taking a specific field as the target field (such as the financial field) as an example, the code generation system 20 can obtain incremental test text code pairs and incremental test methods from the code warehouse or integrated development and testing environment of the target field, and then the code generation system 20 can vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library. By adopting the incremental update mechanism, the text code pair vector library can store the latest test step descriptions and test text code pairs formed by the test code, and the test method vector library can store the latest test methods.

[0114] The code generation system 20 can vectorize the test step description to obtain a test step description vector, and then calculate the distance between the test step description vector and the text code to the vector in the vector library. The distance can be a cosine distance or a Euclidean distance. The distance of the vector can be used to characterize the similarity of the vector. The code generation system 20 can determine similar step descriptions based on the distance of the vector, and then obtain the test code snippet corresponding to the similar step description. Similarly, the code generation system 20 can perform semantic analysis on the test step description to determine the test method whose semantic similarity meets the requirements. Among them, the semantic similarity meets the requirements can be that the semantic similarity is greater than a set value. Based on this, the test method is also called a semantically similar test method. The code generation system 20 can obtain semantically similar test method call samples from the test method vector library. The test method call sample corresponding to the test step description can be the above-mentioned semantically similar test method call sample.

[0115] S606 , the code generation system 20 generates test codes corresponding to the test step description by performing reasoning through LLM according to the test code snippets and test method calling examples corresponding to the similar test step descriptions.

[0116] In specific implementation, the code generation system 20 can also obtain hints from the context of the test step description, for example, obtaining hints from the code file where the test step description is located. Accordingly, the code generation system 20 can splice the hints with the test code snippets and test method call samples corresponding to similar test step descriptions, and then input the test code snippets and test method call samples with the spliced ​​hints into the LLM for reasoning to obtain the test code corresponding to the test step description.

[0117] The prompt may include at least one of global information in the context, use case level information, step level information, or cross-use case context information. Figure 7 As shown, global information (referred to as global info) may include product information, file information, and test language framework information; product information includes product line and product development unit (PDU); file information includes file name, Git repository, and file path information; and use case level information (referred to as use case level info) includes at least one of TC attribute info, import library, use case base class info, and variable initialization definition. Among them, TC attribute info may include at least one of use case class name, use case name, test type, test activity, characteristic, test environment type, precondition, and test step description; import library includes at least one of import class signature, annotation or import method signature, and annotation; and use case base class info includes at least one of base class signature, annotation or base class method signature, and annotation. Step level information (referred to as step level info) includes at least one of step text description, step test script code block, and preceding step code. Cross-use case context information (referred to as cross-use case info) includes at least one of test suite information (test suite info), environment information (referred to as ENVinfo), and configuration information (referred to as Config info). The test suite info includes at least one of the test suite attribute definition, the test suite method signature, and the annotation.

[0118] LLM is a model used to automatically generate test code. The model can be a general LLM, such as L0 LLM, or an LLM trained with training data from the target domain, such as the aforementioned L1 LLM or L2 LLM. L1 LLM and L2 LLM can be trained with a general base model such as L1 LLM. The LLM training process is described in detail below.

[0119] The code generation system 20 can obtain the test engineering code of the target domain, and then train the general base model according to the test engineering code of the target domain to obtain the LLM. The LLM can automatically generate test code for the target domain. The code generation system 20 can obtain the prompt prompt from the context of the test engineering code, use the real code snippet corresponding to the test step description in the test engineering code of the target domain as the answer, construct a prompt-answer pair according to the prompt and the answer, and fine-tune the L0 LLM or L1 LLM through the prompt-answer pair to obtain the L2 LLM.

[0120] In some possible implementations, the code generation system 20 can pre-train a general base model (such as the above-mentioned L0 LLM) according to the test engineering code of the target domain to obtain a test code generation pre-training model (such as the above-mentioned L1 LLM). The test code generation pre-training model can be a pre-training model for the target domain for automatically generating test codes. Then the code generation system 20 can perform supervised fine-tuning on the test code generation pre-training model according to the text code pairs extracted from the test engineering code of the target domain to obtain an LLM (such as the above-mentioned L2 LLM).

[0121] In some other possible implementations, the code generation system 20 may directly perform supervised fine-tuning on a general base model (such as L0 LLM) based on text-code pairs extracted from the test engineering code of the target domain to obtain an LLM (such as L2LLM).

[0122] S608. The code generation system 20 presents the test code corresponding to the test step description to the user.

[0123] Specifically, the code generation system 20 can present the test code corresponding to the test step description to the user through a graphical user interface or a command line interface. Further, the graphical user interface can also include feedback controls, such as an accept control, a modify control, and a reject control. When the user triggers the accept control, it means accepting the test code generated by LLM reasoning. When the user triggers the modify control, the test code generated by LLM reasoning can be modified. When the user triggers the reject control, it means rejecting the test code generated by LLM reasoning.

[0124] In some possible implementations, the code generation system 20 may determine the code retention rate and present the code retention rate to the user to achieve objective evaluation. The code generation system 20 may also present the test code generated by LLM reasoning and the test code modified by the user to the user to achieve subjective evaluation.

[0125] Based on the above description, the code generation method of the present application combines retrieval and LLM reasoning. It first searches the text code pair vector library and the test method vector library to obtain test code snippets and test method call samples corresponding to similar test step descriptions (for example, test method call samples with similar semantics). Based on the test code snippets and test method call samples, LLM reasoning is performed instead of reasoning at the function granularity. This can automatically generate complex test code (such as system test code) and improve the efficiency of automated development of complex test scenarios (such as integration testing and system testing).

[0126] The above is an introduction to the code generation method from the perspective of the code generation system 20. The following is a detailed description of the code generation method of the present application in combination with a specific application scenario.

[0127] In this application scenario, the code generation method may include the following steps:

[0128] Step 1: Standardization and specification checking and cleaning of LLM training corpus data format

[0129] 1. The code generation system 20 determines the standard format of L1 & L2 LLM training corpus data.

[0130] In some examples, a sample of the L2 test text code pair corpus data format is shown below:

[0131]

[0132]

[0133]

[0134] 2. The code generation system 20 uses inspection rules and cleaning tools to check whether the above corpus complies with the specifications according to the specification categories, subcategories, and inspection items, and automatically or manually cleans and rectifies the violations.

[0135] like Figure 8 As shown, for the task of generating script code from system test case text, the valid examples of standard check items include the following categories: uniqueness check, validity check, integrity check, consistency check and security check. Each category can also include several check items. Taking consistency check as an example, consistency check can be subdivided into text code consistency (also known as text script consistency check), case text quality check, test method quality check, test method density check and test method style check. Furthermore, text code consistency check can be further subdivided into the following check items: text similarity, inconsistent test code implementation; similar test code implementation, inconsistent text description.

[0136] Fig. 9It also shows the process of data inspection and cleaning tools. Fig.10 The results of inspection and cleaning according to the above inspection specifications are shown. Fig.10 A sample data quality metrics dashboard is also provided.

[0137] Step 2: The code generation system 20 executes the prompt project and performs model training.

[0138] 1) The code generation system 20 incrementally pre-trains L0 based on the test engineering code corpus data after cleaning in step 1 to obtain L1 LLM.

[0139] It should be noted that the code generation method of the present application may not execute this step. For example, the code generation system 20 may directly train L0 to obtain L2 LLM.

[0140] 2) Prompt project: The code generation system 20 extracts the context of the test case based on the test project code corpus data after cleaning in step 1 and constructs a prompt. The information that can be used to construct the prompt includes but is not limited to the following: Figure 7 The information dimensions shown, for example, can include global information such as product, path, language, test framework, etc.; test case name, purpose, classification, full text of step description, import dependency library, base class info, etc., use case level information; text description of the step to be inferred, reference code for the reasoning step, etc., step level information; and information such as test suites across use cases, environment configuration, etc.

[0141] 3) The code generation system 20 constructs prompt-answer pairs based on the real code snippets corresponding to the test step description as answers, and then fine-tunes the L0 LLM or L1 LLM in a supervised manner to obtain the L2 LLM.

[0142] Step 3: The code generation system 20 generates test code by reasoning based on the newly input test step description (e.g., test case text) and the prompts in the context. The framework of a feasible reasoning tool (e.g., reasoning subsystem 202) is as follows: Fig.11 As shown, the reasoning subsystem 202 realizes reasoning enhancement by retrieving the latest local text code vector library and test method vector library. Among them, the reasoning subsystem 202 includes a front-end plug-in and an API interface for interacting with the user, obtaining the test step description and context, and routing the task request information including the test step description and context to the corresponding agent service after the Agent scheduling system is limited and authenticated, completing the context extraction and vector library similar information retrieval, and then splicing the prompt prompt to request the LLM reasoning service to generate the test code by reasoning. The test code can be post-processed and then returned to the API interface through the action module, and finally returned to the user for confirmation through the IDE plug-in.

[0143] It should be noted that the code generation system 20 may include multiple reasoning scenarios when reasoning about the test code, which are described below respectively.

[0144] Reasoning scenario 1: step-level test code generation. Specifically, the code generation system 20 receives the step to be reasoned input by the user, generates a step-level script code snippet based on the step-level text description and its context, and returns it.

[0145] Reasoning scenario 2: Use case level test code generation. Specifically, the code generation system 20 receives all steps and contexts of the use case to be reasoned input by the user, generates test codes for all steps in batches and returns them.

[0146] Step 4: The code generation system 20 performs result evaluation. The detailed steps of result evaluation are as follows: Fig.12 As shown, first, the code generation system 20 creates an evaluation task in response to a task configuration operation triggered by a user. The task configuration operation includes selecting the version of the model to be tested, the task type, the evaluation data set, and the evaluation indicators. Secondly, the code generation system 20 performs verification set reasoning. Next, the code generation system 20 compares the test code generated by reasoning with the actual test code in the verification set, calculates the index value of the objective evaluation index, and presents the index value of the objective evaluation index and the details of the reasoning results to the user for further subjective evaluation. The calculation formula and sample of the objective evaluation index are as follows: Fig.13 As shown, no further description is given here.

[0147] Based on the above description, the code generation method provided by the present application performs enhanced reasoning on the vector library and the test method vector library by retrieving the latest local text code, thereby solving the problem that the newly added text code pairs cannot be fed to L2Fine-Tune immediately, and solving the problem of using the newly added test methods to assist in generating test codes. Moreover, the method provides a set of inspection specifications for checking the training corpus, which can automatically check the corpus quantitatively, continuously iterate the management corpus, and significantly improve the model training effect. When training the model, the method improves the test code generation effect for one-party, two-party, and three-party libraries in specific fields through pre-training and supervised fine-tuning. The present application also provides a set of metrics for evaluating the test code generation effect, so as to achieve quantitative and objective evaluation of the effect of the task of generating code from test text.

[0148] Based on the aforementioned code generation method, the present application also provides a code generation system. Fig.14 As shown, the code generation system 20 includes:

[0149] Interaction module 1402, used to receive a test step description input by a user;

[0150] A retrieval module 1404 is used to retrieve a text code pair vector library and a test method vector library according to the test step description to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes a description text and a vector formed by the corresponding test code, and the test method vector library includes a test method vector;

[0151] Reasoning module 1406, used to generate test code corresponding to the test step description by reasoning through a large language model LLM according to the test code snippets and test method call examples corresponding to the similar test step description;

[0152] The interaction module 1402 is further configured to present the test code corresponding to the test step description to the user.

[0153] The interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may be the modules in the aforementioned reasoning subsystem 202. For example, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may be implemented by hardware or by software.

[0154] When implemented by software, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 can be applications running on a computing device, such as a computing engine, etc. Among them, the application can be provided in the form of a virtualization service. Virtualization services may include virtual machine (VM) services, bare metal server (BMS) services, and container services. Among them, the VM service can be a service that virtualizes a virtual machine (VM) resource pool on multiple physical hosts through virtualization technology to provide users with VMs on demand for use. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide users with BMS on demand for use. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide users with containers on demand for use. VM is a simulated virtual computer, that is, a logical computer. BMS is a high-performance computing service that can be elastically scalable, and its computing performance is no different from that of a traditional physical machine, and it has the characteristics of secure physical isolation. Containers are a kernel virtualization technology that can provide lightweight virtualization to achieve the purpose of isolating user space, processes, and resources. It should be understood that the VM service, BMS service and container service in the above-mentioned virtualization services are only specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.

[0155] When implemented by hardware, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may include at least one computing device, such as a server, etc. Alternatively, the interaction module 1402, the retrieval module 1404, and the reasoning module 1406 may also be implemented by using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0156] In some possible implementations, the system 20 further includes:

[0157] The vector library management module 1408 is used to obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library.

[0158] The vector library management module 1408 may be a module in the reasoning subsystem 202, such as a software module or a hardware module in the reasoning subsystem 202, and this embodiment does not limit this. When implemented by software, the vector library management module 1408 may be an application running on a computing device, and the application may be provided in the form of a virtualized service, such as a BMS service, a VM service, or a container service. When implemented by hardware, the vector library management module 1408 may include at least one computing device, such as a server, etc. Alternatively, the vector library management module 1408 may also be a device implemented using an application-specific integrated circuit ASIC, or a programmable logic device PLD, etc.

[0159] In some possible implementations, the interaction module 1402 is further configured to:

[0160] Taking hints from the context of the test step description;

[0161] The reasoning module 1406 is specifically used for:

[0162] splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions;

[0163] The test code snippet spliced ​​with the prompt and the test method calling sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.

[0164] In some possible implementations, the prompt includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.

[0165] In some possible implementations, the system 20 further includes:

[0166] The training module 1410 is used to obtain the test engineering code of the target domain, train the universal base model according to the test engineering code of the target domain, and obtain the LLM.

[0167] Among them, the training module 1410 may be a module in the training subsystem 206. The training module 1410 may be implemented by software or by hardware. When implemented by software, the training module 1410 may be an application running on a computing device, which may be provided in the form of a virtualized service, such as a BMS service, a VM service, or a container service. When implemented by hardware, the training module 1410 may include at least one computing device, such as a server, etc. Alternatively, the training module 1410 may also be a device implemented using an application-specific integrated circuit ASIC, or a programmable logic device PLD, etc.

[0168] In some possible implementations, the training module 1410 is specifically used to:

[0169] Pre-training a universal base model according to the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model according to text code pairs extracted from the test engineering code of the target domain to obtain the LLM; or

[0170] The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.

[0171] In some possible implementations, the system 20 further includes:

[0172] The evaluation module 1412 is used to obtain a verification set, the verification set includes text code pairs of the target domain, the text code pairs include description texts and real code snippets, use the LLM model to infer the description texts in the verification set to obtain generated code snippets, and determine the index value of the evaluation index according to the generated code snippets and the real code snippets, the evaluation index includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate, and test parameter assignment precision rate;

[0173] The interaction module 1402 is further configured to present the indicator value of the evaluation indicator to the user.

[0174] The evaluation module 1412 may be a module in the training subsystem 206. The evaluation module 1412 may be implemented by software or by hardware. When implemented by software, the evaluation module 1412 may be an application running on a computing device, which may be provided in the form of a virtualized service, such as a BMS service, a VM service, or a container service. When implemented by hardware, the evaluation module 1412 may include at least one computing device, such as a server. Alternatively, the evaluation module 1412 may also be a device implemented by an application specific integrated circuit ASIC or a programmable logic device PLD.

[0175] In some possible implementations, the interaction module 1402 is further configured to:

[0176] The generated code snippet and the actual code snippet are presented to the user.

[0177] The present application also provides a computing device 1500. Fig.15 As shown, computing device 1500 includes: bus 1502, processor 1504, memory 1506 and communication interface 1508. Processor 1504, memory 1506 and communication interface 1508 communicate through bus 1502. Computing device 1500 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 1500.

[0178] The bus 1502 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.15The bus 1502 may include a path for transmitting information between various components of the computing device 1500 (eg, the memory 1506, the processor 1504, and the communication interface 1508).

[0179] The processor 1504 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0180] The memory 1506 may include a volatile memory, such as a random access memory (RAM). The memory 1506 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). The memory 1506 stores executable program code, and the processor 1504 executes the executable program code to implement the aforementioned code generation method. Specifically, the memory 1506 stores instructions for the code generation system 20 to execute the code generation method.

[0181] The communication interface 1508 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1500 and other devices or communication networks.

[0182] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0183] like Fig.16 As shown, the computing device cluster includes at least one computing device 1500. The memory 1506 in one or more computing devices 1500 in the computing device cluster may store the same code generation system 20 for executing instructions of the code generation method.

[0184] In some possible implementations, one or more computing devices 1500 in the computing device cluster may also be used to execute some instructions of the code generation system 20 for executing the code generation method. In other words, a combination of one or more computing devices 1500 may jointly execute the instructions of the code generation system 20 for executing the code generation method.

[0185] It should be noted that the memory 1506 in different computing devices 1500 in the computing device cluster may store different instructions for executing partial functions of the code generation system.

[0186] Fig.17 A possible implementation is shown. Fig.17 As shown, two computing devices 1500A and 1500B are connected via a communication interface 1508. The memory in the computing device 1500A stores instructions for executing the functions of the interaction module 1402 and the retrieval module 1404. The memory in the computing device 1500B stores instructions for executing the functions of the reasoning module 1406. In other words, the memory 1506 of the computing devices 1500A and 1500B jointly stores instructions for the code generation system 20 to execute the code generation method. Further, the computing device 1500A, the computing device 1500B or other computing devices may also store instructions for executing the functions of the vector library management module 1408, the training module 1410, and the evaluation module 1412.

[0187] Fig.17 The connection mode between the computing device clusters shown may be considered to be that the code generation method provided by the present application requires a lot of computing power for reasoning. Therefore, it is considered to hand over the functions implemented by the reasoning module 1406 to the computing device 1500B for execution.

[0188] It should be understood that Fig.17 The functionality of the computing device 1500A shown in FIG. 1 may also be implemented by multiple computing devices 1500. Similarly, the functionality of the computing device 1500B may also be implemented by multiple computing devices 1500.

[0189] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.18 A possible implementation is shown. Fig.18As shown, two computing devices 1500C and 1500D are connected via a network. Specifically, they are connected to the network via the communication interface in each computing device. In this type of possible implementation, the memory 1506 in the computing device 1500C stores instructions for executing the functions of the interaction module 1402 and the retrieval module 1404. At the same time, the memory 1506 in the computing device 1500D stores instructions for executing the functions of the reasoning module 1406. Further, the computing device 1500C, the computing device 1500D or other computing devices may also store instructions for executing the functions of the vector library management module 1408, the training module 1410, and the evaluation module 1412.

[0190] Fig.18 The connection method between the computing device clusters shown may be that considering that the code generation method provided in the present application requires a large amount of computing power for reasoning, it is considered that the functions implemented by the reasoning module 1406 are handed over to the computing device 1500D for execution.

[0191] It should be understood that Fig.18 The functions of the computing device 1500C shown in FIG. 1500A may also be completed by multiple computing devices 1500. Similarly, the functions of the computing device 1500D may also be completed by multiple computing devices 1500.

[0192] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned application to the code generation system 20 for executing the code generation method.

[0193] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the above-mentioned code generation method.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A code generation method, It is characterized in that The method is performed by a code generation system, wherein the code generation system is used to automatically generate test code, and the method includes: Receive test step descriptions from user input; According to the test step description, a text code pair vector library and a test method vector library are retrieved to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes a description text and a vector formed by the corresponding test code, and the test method vector library includes a test method vector; According to the test code snippets and test method call examples corresponding to the similar test step descriptions, reasoning is performed through the large language model LLM to generate the test code corresponding to the test step description; The test code corresponding to the test step description is presented to the user.

2. The method according to claim 1, It is characterized in that The method further comprises: Obtain incremental test text code pairs and incremental test methods from the target domain's code repository or integrated development and testing environment; The incremental test text code pairs are vectorized and added into the text code pair vector library, and the incremental test methods are vectorized and added into the test method vector library.

3. The method according to claim 1 or 2, It is characterized in that The method further comprises: Taking hints from the context of the test step description; The method of performing reasoning based on the test code snippets and test method call examples corresponding to the similar test step descriptions through a large language model (LLM) to obtain the test code corresponding to the test step descriptions includes: splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions; The test code snippet spliced ​​with the prompt and the test method calling sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.

4. The method according to claim 3, It is characterized in that The hint includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.

5. The method according to any one of claims 1 to 4, It is characterized in that The method further comprises: Obtain the test engineering code of the target domain; The LLM is obtained by training a general base model according to the test engineering code of the target domain.

6. The method according to claim 5, It is characterized in that The step of training a base model according to the test engineering code of the target domain to obtain the LLM comprises: Pre-training a universal base model according to the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model according to text code pairs extracted from the test engineering code of the target domain to obtain the LLM; or The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.

7. The method according to claim 5 or 6, It is characterized in that The method further comprises: Acquire a verification set, wherein the verification set includes text-code pairs of the target domain, and the text-code pairs include description text and real code snippets; Use the LLM model to infer the description text in the verification set to obtain a generated code snippet; Determine an index value of an evaluation index according to the generated code snippet and the real code snippet, wherein the evaluation index includes at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter assignment recall rate, and test parameter assignment precision rate; The indicator value of the evaluation indicator is presented to the user.

8. The method according to claim 7, It is characterized in that The method further comprises: The generated code snippet and the actual code snippet are presented to the user.

9. The method according to any one of claims 1 to 8, It is characterized in that The test step description includes a description based on a natural language or a description based on a domain specific language.

10. A code generation system, It is characterized in that The code generation system is used to automatically generate test code, and the system includes: An interactive module, used for receiving a test step description input by a user; A retrieval module is used to retrieve a text code pair vector library and a test method vector library according to the test step description to obtain test code snippets and test method call samples corresponding to similar test step descriptions, wherein the text code pair vector library includes a description text and a vector formed by the corresponding test code, and the test method vector library includes a test method vector; An inference module, configured to generate test code corresponding to the test step description by performing inference through a large language model (LLM) based on the test code snippets and test method call examples corresponding to the similar test step descriptions; The interaction module is further used to present the test code corresponding to the test step description to the user.

11. The system according to claim 10, It is characterized in that The system further comprises: The vector library management module is used to obtain incremental test text code pairs and incremental test methods from the code repository or integrated development and testing environment of the target domain, vectorize the incremental test text code pairs and add them to the text code pair vector library, and vectorize the incremental test methods and add them to the test method vector library.

12. The system according to claim 10 or 11, It is characterized in that The interaction module is also used for: Taking hints from the context of the test step description; The reasoning module is specifically used for: splicing the test code snippets and test method call samples corresponding to the prompts and the similar test step descriptions; The test code snippet spliced ​​with the prompt and the test method calling sample are input into the LLM for reasoning to obtain the test code corresponding to the test step description.

13. The system according to claim 12, It is characterized in that The hint includes at least one of global information, use case level information, step level information, or cross-use case context information in the context.

14. A system according to any one of claims 10 to 13, It is characterized in that The system further comprises: The training module is used to obtain the test engineering code of the target field, train the universal base model according to the test engineering code of the target field, and obtain the LLM.

15. The system according to claim 14, It is characterized in that The training module is specifically used for: Pre-training a universal base model according to the test engineering code of the target domain to obtain a test code generation pre-training model, and performing supervised fine-tuning on the test code generation pre-training model according to text code pairs extracted from the test engineering code of the target domain to obtain the LLM; or, The LLM is obtained by performing supervised fine-tuning on a general base model based on text-code pairs extracted from test engineering codes in the target domain.

16. The system according to claim 14 or 15, It is characterized in that The system further comprises: An evaluation module is used to obtain a verification set, wherein the verification set includes text-code pairs of the target domain, wherein the text-code pairs include description texts and real code snippets, and use the LLM model to infer the description texts in the verification set to obtain generated code snippets, and determine the index values ​​of evaluation indicators according to the generated code snippets and the real code snippets, wherein the evaluation indicators include at least one of code similarity, code retention rate, test method recall rate, test method precision rate, test parameter recall rate, test parameter precision rate, test parameter value recall rate, and test parameter value precision rate; The interaction module is further used to present the indicator value of the evaluation indicator to the user.

17. The system according to claim 16, It is characterized in that The interaction module is also used for: The generated code snippet and the actual code snippet are presented to the user.

18. A system according to any one of claims 10 to 17, It is characterized in that The test step description includes a description based on a natural language or a description based on a domain specific language.

19. A computing device cluster, It is characterized in that The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the computing device cluster executes the code generation method according to any one of claims 1 to 9.

20. A computer-readable storage medium, It is characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.

21. A computer program product, It is characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Code generation method and device, storage medium and electronic equipment

    CN120447880A

  • Code generation method and device, storage medium and electronic device

    CN120447880B

  • Multi-Agent software code automatic generation system and method

    CN121143765A