Unit test code generation method and device, equipment and storage medium
By using a large language model to query relevant class names from source code metadata and generate unit test codes, the problem of time-consuming and labor-intensive generation of unit test codes in the existing technology is solved, and efficient and accurate unit test code generation is achieved.
Patent Information
- Application Number
- CN202411979360.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, unit test code generation is time-consuming and labor-intensive, and it is difficult to ensure accuracy.
By obtaining the program source code of the target program and using a large language model to query the class names related to the target unit test code from the source code metadata, the corresponding unit test code is generated.
It improves the efficiency and accuracy of unit test code generation, reduces the amount of data processed, ensures the integrity and correctness of the code, and realizes automatic generation.
Smart Images

Figure CN119938000A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software processing technology, and in particular to a unit test code generation method, device, equipment and storage medium. Background Art
[0002] Unit test code refers to automated test code that verifies the smallest testable unit in a program, and is usually used to verify the correctness of the smallest testable unit. The smallest testable unit includes but is not limited to functions, methods or classes in the program source code.
[0003] In the related art, unit test codes are often written by developers based on program source codes, which is time-consuming and labor-intensive, and it is difficult to ensure the accuracy of the unit test codes. Summary of the invention
[0004] The embodiments of the present invention provide a unit test code generation method, apparatus, device and storage medium, which are used to improve the efficiency and accuracy of unit test code generation.
[0005] In a first aspect, an embodiment of the present invention provides a method for generating unit test code, the method comprising:
[0006] In response to a unit test code generation request triggered by a user, obtaining a program source code of a target program, wherein the unit test code generation request is used to request generation of a target unit test code corresponding to a first program source code in the program source code;
[0007] Inputting the program source code and the first program source code into a large language model, so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code;
[0008] Determining source code metadata corresponding to the program source code, wherein the amount of data corresponding to the source code metadata is smaller than the amount of data corresponding to the program source code;
[0009] Determining, from the source metadata, target source metadata that matches the class name of the at least one class;
[0010] A second program source code corresponding to the target source code metadata is input into the large language model so that the large language model generates a target unit test code corresponding to the first program source code, wherein the program source code includes the second program source code.
[0011] In a second aspect, an embodiment of the present invention provides a unit test code generation device, the device comprising:
[0012] An acquisition module, configured to acquire a program source code of a target program in response to a unit test code generation request triggered by a user, wherein the unit test code generation request is used to request generation of a target unit test code corresponding to a first program source code in the program source code;
[0013] a processing module, configured to input the program source code and the first program source code into a large language model, so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code; determine source code metadata corresponding to the program source code, wherein the amount of data corresponding to the source code metadata is smaller than the amount of data corresponding to the program source code; and determine target source code metadata matching the class name of the at least one class from the source code metadata;
[0014] A generation module is used to input a second program source code corresponding to the target source code metadata into the large language model so that the large language model generates a target unit test code corresponding to the first program source code, and the program source code includes the second program source code.
[0015] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the unit test code generation method described in the first aspect.
[0016] In a fourth aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor can at least implement the unit test code generation method described in the first aspect.
[0017] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising: a computer program, when the computer program is executed by a processor of an electronic device, the processor can at least implement the unit test code generation method as described in the first aspect.
[0018] In the solution provided by an embodiment of the present invention, in response to a unit test code generation request triggered by a user, first, the program source code of the target program is obtained, wherein the unit test code generation request is used to request the generation of a target unit test code corresponding to a first program source code in the program source code. Afterwards, the program source code and the first program source code are input into a large language model so that the large language model outputs the class name of at least one class used by the method corresponding to the first program source code. Next, the source code metadata corresponding to the program source code is determined, and the amount of data corresponding to the source code metadata is less than the amount of data corresponding to the program source code. Finally, the target source code metadata that matches the class name of at least one class is determined from the source code metadata; the second program source code corresponding to the target source code metadata is input into the large language model so that the large language model generates the target unit test code corresponding to the first program source code, wherein the program source code includes the second program source code.
[0019] In this solution, when generating the target unit test code corresponding to the first program source code, the second program source code used to generate the target unit test code is obtained by querying the source code metadata corresponding to the program source code of the target program for the target metadata that matches at least one class used by the method corresponding to the first program source code. Compared with directly obtaining the second program source code from the program source code, the amount of data processed is reduced, which can effectively improve the generation efficiency of the target unit test code; in addition, compared with the traditional solution of splitting the program source code and then obtaining the second program source code from the split program source code blocks, the second program source code is obtained based on the source code metadata, which can ensure the integrity of the obtained second program source code and improve the accuracy of the target unit test code generation. In addition, in this solution, combined with a large language model, the automatic generation of the target unit test code can also be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A flowchart of a method for generating unit test code provided by an embodiment of the present invention;
[0022] Figure 2 A flowchart of another unit test code generation method provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of a unit test code generation process provided by an embodiment of the present invention;
[0024] Figure 4 A flowchart of a method for determining target source code metadata provided by an embodiment of the present invention;
[0025] Figure 5 A schematic diagram of the structure of a unit test code generation device provided by an embodiment of the present invention;
[0026] Figure 6 For Figure 5 A schematic structural diagram of an electronic device corresponding to the unit test code generating device provided in the illustrated embodiment. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0029] Some embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the case where there is no conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.
[0030] For ease of understanding, relevant concepts involved in the embodiments of the present invention are described below.
[0031] Program source code refers to a set of instructions written by programmers in a high-level programming language (such as C, C++, Java, Python, etc.) to implement a specific application.
[0032] Unit test code refers to automated test code that verifies the smallest testable unit in the program source code, and is usually used to verify the correctness of the smallest testable unit. The smallest testable unit refers to a code snippet in the program source code that has clear input and output, relatively independent functions, and can be verified separately, including but not limited to functions, methods, or classes in the program source code.
[0033] Source code metadata is used to describe the structured data of the program source code. The amount of data corresponding to the source code metadata is smaller than the amount of data of the program source code.
[0034] Large Language Model (LLM) refers to a large-scale parameter language model learned using a large-scale data corpus based on a deep learning framework. It can be used to process a variety of natural language tasks such as text classification, question and answer, and dialogue.
[0035] In actual applications, to generate unit test code for a minimum testable unit, it is necessary to first obtain all program source codes related to the minimum testable unit from the program source code, that is, all the code required for the code snippet corresponding to the minimum testable unit to run successfully; then, based on the program source code related to the minimum testable unit obtained from the program source code, generate the unit test code corresponding to the minimum testable unit.
[0036] In the related art, when it is necessary to generate unit test code corresponding to a minimum testable unit, the program source code corresponding to the program (also called a project) is usually first divided into multiple different program source code blocks according to a specific size, for example, according to the division rule that each program source code block contains 30 instructions, the program source code is divided into multiple different program source code blocks, etc. Afterwards, the program source code related to the minimum testable unit is obtained from these program source code blocks to generate the unit test code corresponding to the minimum testable unit.
[0037] In actual applications, the program source code corresponding to the program is stored in different source code files according to classes. One class corresponds to one source code file. Usually, the name of the source code file is the name of the class. One source code file contains several instructions for constituting the program source code. The number of instructions contained in the source code files corresponding to different classes can be the same or different. It should be emphasized that the class mentioned in the embodiments of the present invention refers to a data type in programming, not the type usually described.
[0038] Based on the storage method of program source code, when generating unit test code corresponding to a minimum testable unit according to the unit test code generation method in the related art, the code instructions in the source code file corresponding to the same class may be divided into different program source code blocks, which will make it difficult to obtain the complete content when obtaining the program source code related to the minimum testable unit from the program source code. For example, the program source code corresponding to a class is divided into different program source code blocks. When obtaining the program source code related to the minimum testable unit, only the partial program source code corresponding to the class contained in the partial program source code blocks is obtained. When the obtained program source code related to the minimum testable unit is incomplete, it will affect the accuracy of the generated unit test code.
[0039] In addition, the program source code usually corresponds to a large amount of data. In the application scenario of generating unit test code, the program source code is directly stored, and the program source code related to the minimum testable unit is obtained from the stored program code. On the one hand, it will cause a large amount of storage space to be occupied, and on the other hand, it will cause greater data processing pressure.
[0040] To solve at least one of the above technical problems, an embodiment of the present invention provides a unit test code generation method: in response to a unit test code generation request triggered by a user, a program source code of a target program is obtained, the unit test code generation request is used to request the generation of a target unit test code corresponding to a first program source code in the program source code; the program source code and the first program source code are input into a large language model, so that the large language model outputs the class name of at least one class used by the method corresponding to the first program source code; source code metadata corresponding to the program source code is determined, the amount of data corresponding to the source code metadata is less than the amount of data corresponding to the program source code; target source code metadata matching the class name of at least one class is determined from the source code metadata; a second program source code corresponding to the target source code metadata is input into the large language model, so that the large language model generates a target unit test code corresponding to the first program source code, the program source code includes the second program source code.
[0041] Among them, when generating the target unit test code corresponding to the first program source code, the second program source code used to generate the target unit test code is obtained by querying the source code metadata corresponding to the program source code of the target program for the target metadata that matches at least one class used by the method corresponding to the first program source code. Compared with directly obtaining the second program source code from the program source code, the amount of data processed is reduced, which can effectively improve the generation efficiency of the target unit test code; in addition, compared with the traditional solution of splitting the program source code and then obtaining the second program source code from the split program source code blocks, the second program source code is obtained based on the source code metadata, which can ensure the integrity of the obtained second program source code and improve the accuracy of the target unit test code generation. In addition, in this solution, combined with a large language model, the automatic generation of the target unit test code can also be achieved.
[0042] The unit test code generation method provided by the embodiment of the present invention is described in detail below in conjunction with specific embodiments.
[0043] The unit test code generation method provided in the embodiment of the present invention can be executed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, or a server. The server can be a physical server including an independent host, or a virtual server, or a cloud server or a server cluster.
[0044] Figure 1 A flowchart of a method for generating unit test code provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the following steps may be included:
[0045] 101. In response to a unit test code generation request triggered by a user, a program source code of a target program is obtained, wherein the unit test code generation request is used to request generation of a target unit test code corresponding to a first program source code in the program source code.
[0046] The first program source code is a code snippet corresponding to any minimum testable unit in the source code program. The program source code of the target program is obtained according to the program source code path specified by the user when triggering a unit test code generation request.
[0047] 102. Input the program source code and the first program source code into a large language model, so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code.
[0048] Among them, the large language model has learned a large amount of program source code during the training phase, and has the ability to understand the code structure and parse out the class name of at least one class used by the method corresponding to a certain code fragment (ie, the first program source code) from the program source code.
[0049] As described above, program source codes are stored in different source code files according to classes, and the name of the source code file is the name of the class. Therefore, in this embodiment, the second program source code related to the first source code program and used to generate the target unit test code is obtained from the program source code through the class name of at least one class used by the method corresponding to the first source code program. The second program source code is all the code required for the first program source code to run successfully.
[0050] 103. Determine source code metadata corresponding to the program source code, and the amount of data corresponding to the source code metadata is smaller than the amount of data corresponding to the program source code.
[0051] 104. Determine, from the source metadata, target source metadata that matches the class name of at least one class.
[0052] In an embodiment of the present invention, in order to reduce the amount of data required to be processed when obtaining the second program source code and save storage space, after obtaining the program source code, the program source code is first parsed and information is extracted to determine the source code metadata corresponding to the program source code, and the source code metadata is stored; then, the second program source code is obtained based on the stored source code metadata.
[0053] Among them, since the source code metadata is the result of information extraction of the program source code, the amount of data corresponding to the source code metadata is smaller than the amount of data corresponding to the program source code, the storage space occupied by storing the source code metadata is smaller than the storage space occupied by storing the program source code, and the amount of data to be processed when obtaining the second program source code based on the source code metadata is smaller than the amount of data to be processed when obtaining the second program source code directly from the program source code.
[0054] During the specific implementation process, the data category included in the source code metadata can be pre-defined based on the information required to obtain the second program source code; thereby, data matching the data category can be obtained from the program source code to generate source code metadata corresponding to the data source code.
[0055] In the embodiment of the present invention, the second program source code related to the first source code program and used to generate the target unit test code is obtained from the program source code through the class name of at least one class used by the method corresponding to the first source code program. Therefore, optionally, the data category included in the source code metadata can be: the class name of the target class. The target class includes but is not limited to: parent class, parent interface, inner class, custom class, custom enumeration, etc.
[0056] It is understandable that there is a corresponding relationship between the program source code and the source code metadata, that is, the corresponding code snippet in the program source code can be determined based on a certain source code metadata. Therefore, the target source code metadata matching the class name of the at least one class used by the method corresponding to the first program source code and the class name of the target class included in the source code metadata can be determined from the source code metadata; then, based on the corresponding relationship between the program source code and the source code metadata, the program source code corresponding to the target source code metadata is determined to be the second program source code.
[0057] 105. Input a second program source code corresponding to the target source code metadata into the large language model, so that the large language model generates a target unit test code corresponding to the first program source code, wherein the program source code of the target program includes the second program source code.
[0058] Specifically, the second program source code corresponding to the target source code metadata can be first filled into a preset slot in a preset prompt word template to generate a prompt word for generating the target unit test code, wherein the prompt word template includes description information of the unit test code generation task. Afterwards, the prompt word is input into the large language model, so that the large language model generates the target unit test code corresponding to the first program source code under the guidance of the prompt word.
[0059] In summary, in an embodiment of the present invention, when generating a target unit test code corresponding to a first program source code, a second program source code for generating a target unit test code is obtained by querying the source code metadata corresponding to the program source code of the target program for target metadata that matches at least one class used by the method corresponding to the first program source code. Compared with directly obtaining the second program source code from the program source code, the amount of data processed is reduced, and the generation efficiency of the target unit test code can be effectively improved. In addition, obtaining the second program source code based on the source code metadata is compared to the traditional solution of splitting the program source code and then obtaining the second program source code from the split program source code blocks. The second program source code is obtained based on the source code metadata, which can ensure the integrity of the obtained second program source code and improve the accuracy of the target unit test code generation. In addition, in this solution, combined with a large language model, the automatic generation of the target unit test code can also be achieved.
[0060] The above describes the unit test code generation method provided by the embodiment of the present invention from the dimension of the overall program source code of the target program. The following describes the unit test code generation method provided by the embodiment of the present invention from the dimension of multiple source code files constituting the program source code.
[0061] Figure 2 A flowchart of another method for generating unit test code provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the following steps may be included:
[0062] 201. In response to a unit test code generation request triggered by a user, multiple source code files of a target program are obtained, and the program source code of the target program is determined based on the program source codes respectively contained in the multiple source code files. The unit test code generation request is used to request the generation of a target unit test code corresponding to a first program source code in the program source code.
[0063] 202. Input the program source code and the first program source code into a large language model, so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code.
[0064] 203. Determine multiple source code metadata corresponding to multiple source code files.
[0065] 204. Determine target source code metadata that matches the class name of at least one class from the plurality of source code metadata, and the program source code included in the source code file corresponding to the target source code metadata is the second program source code.
[0066] 205. Input the second program source code into the large language model so that the large language model generates a target unit test code corresponding to the first program source code.
[0067] The specific implementation process of step 202 and step 205 may refer to the aforementioned embodiment and will not be described in detail in this embodiment.
[0068] For ease of understanding, the following combination Figure 3 right Figure 2 The unit test code generation method provided in the embodiment is described.
[0069] Figure 3 A schematic diagram of a unit test code generation process provided by an embodiment of the present invention, such as Figure 3 As shown, assuming that according to the program source code path specified by the user when triggering the unit test code generation request, the source code file 1, source code file 2, ..., source code file m (m is a positive integer) corresponding to the target program are obtained, then the program source code of the target program is determined to include: the program source code contained in each source code file from source code file 1 to source code file m.
[0070] In an optional embodiment, in order to conveniently determine the target source code metadata that matches the class name of at least one class from the source code metadata, when determining the source code metadata corresponding to the program source code, the source code metadata corresponding to each source code file is determined based on the source code file; and the multiple source code metadata corresponding to the multiple source code files of the target program are used as the source code metadata corresponding to the program source code of the target program. Among them, generating source code metadata based on the source code file as a unit can effectively avoid the information in the same source code file from being split, and further, can ensure the integrity of the second program source code obtained for generating the target unit test code corresponding to the first program source code.
[0071] like Figure 3 As shown, source code metadata 1 corresponding to source code file 1, source code metadata 2 corresponding to source code file 2, ..., and source code metadata m corresponding to source code file m can be generated. The source code metadata corresponding to the program source code of the target program includes: source code metadata 1, source code metadata 2, ..., and source code metadata m.
[0072] The following takes the target source code file as an example to describe the process of determining the source code metadata corresponding to the target source code file, wherein the target source code file is any one of the source code files corresponding to the target program.
[0073] As mentioned above, the file name of the source code file is consistent with the class name of the class described by the program source code in the source code file. Therefore, when the class name of a class is known, the source code file corresponding to the known class name can be determined by matching the class name with the file name of the source code file, and the program source code contained in the source code file can be obtained.
[0074] Based on this, in this embodiment, the source code metadata corresponding to the target source code file may include the file name of the target source code file. Thus, based on the class name of at least one class used by the method corresponding to the first program source code, the target source code metadata matching the class name of the at least one class may be determined, wherein the file name included in the target source code metadata is consistent with the class name of the at least one class.
[0075] In an embodiment of the present invention, the core purpose of determining the source code metadata corresponding to the program source code of the target program is to determine the second program source code used to generate the target unit test code corresponding to the first program source code while occupying a small amount of storage space. To achieve this purpose, the source code metadata corresponding to the target source code file in this embodiment also includes: the file address of the target source code file. Therefore, after determining the target source code metadata, the source code file corresponding to the target source code metadata can be obtained according to the file address contained in the target source code metadata, and the program source code contained in the source code file is the second program source code used to generate the target unit test code corresponding to the first program source code. By storing the file address in the source code metadata, it is possible to ensure that the second program source code is obtained and reduce the storage space occupied.
[0076] In practical applications, the implementation of a certain class may also depend on other classes. In simple terms, the normal operation of the program source code corresponding to a certain class is inseparable from the program source code of other classes. Based on this, the source code metadata of the target source code file provided by this embodiment also includes the class names of classes that can describe the dependency relationship between classes, such as parent classes, parent interfaces, custom classes, and custom enumerations. These classes describe other source code files that the target source code file depends on through class names, or in other words, describe other classes that the class corresponding to the target source code file depends on. Based on the class names contained in the source code metadata for describing the dependency relationship between source code files, after determining the source code file corresponding to at least one class used by the method corresponding to the first program source code, the other source code files that the source code file depends on can be further determined, thereby effectively ensuring the integrity of the acquired second program source code.
[0077] It should be emphasized that, in the embodiment of the present invention, the class names of the parent class, parent interface, custom class and custom enumeration included in the source code metadata are an example of describing the dependency relationship between different source code files. In practical applications, other information that can describe the dependency relationship between source code files can also be configured in the source code metadata, and is not limited to the class names of the parent class, parent interface, custom class and custom enumeration exemplified in this embodiment.
[0078] In addition, in some cases, other classes that a class depends on may be inner classes. Among them, inner classes refer to classes that are further defined in a class. Usually, the file name of a source code file cannot reflect the inner class. For example, a class AA contains an inner class aa, and the file name of the source code file corresponding to class AA is AA. When the program source code corresponding to the inner class aa is needed, it is usually necessary to first determine the class AA to which the inner class aa belongs, and further, determine the source code file with the file name AA based on class AA, and then obtain the program source code corresponding to the inner class aa from the source code file with the file name AA. Based on this, the source code metadata of the target source code file provided in this embodiment also includes the class name of the inner class.
[0079] Based on the above description, in the specific implementation process, for any target source code file among the multiple source code files of the target program, in the process of generating the source code metadata corresponding to the target source code file, the file name and file address of the target source code file and the class name of the target class defined in the target source code file can be parsed first, wherein the target class includes at least one of a parent class, a parent interface, an inner class, a custom class, and a custom enumeration. Then, the source code metadata corresponding to the target source code file is determined according to the parsed file name, file address, and class name of the target class.
[0080] For example, assuming that the file name of the target source code file is AA, the file address is / address / AA.java, and the target class contained in the target source code file includes the parent class CC and the inner class aa, then the source code metadata corresponding to the target source code file can be: {file name: AA; file address: / address / AA.java; parent class: CC; inner class: aa}.
[0081] The above describes the process of generating source code metadata corresponding to any source code file. Next, the process of determining the target source code metadata that matches the class name of at least one class used by the method corresponding to the first program source code from multiple source code metadata corresponding to the program source code of the target program is described.
[0082] Figure 4 A flowchart of a method for determining target source code metadata provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the following steps may be included:
[0083] 401. Determine first target source code metadata from multiple source code metadata according to the class name of at least one class used by the method corresponding to the first program source code and the file name and class name of the target class respectively included in multiple source code metadata.
[0084] 402. If the first target source code metadata includes the class name of the parent class, parent interface, custom class and / or custom enumeration, then according to the class name of the parent class, parent interface, custom class and / or custom enumeration included in the first target source code metadata, and the file name and class name of the target class respectively included in the multiple source code metadata, the second target source code metadata is determined from the multiple source code metadata, and the second target source code metadata is the source code metadata of the source code files corresponding to other classes on which the at least one class depends.
[0085] 403. Use the first target source code metadata and the second target source code metadata as target source code metadata that matches the class name of at least one class used by the method corresponding to the first program source code.
[0086] In summary, first, based on the class name of at least one class used by the method corresponding to the first program source code, determine from multiple source code metadata the first target source code metadata whose file name is consistent with the class name of the at least one class. Then, determine whether the source code file corresponding to the first target source code metadata depends on other source code files, that is, whether the aforementioned at least one class depends on other classes. If it is determined that the aforementioned at least one class depends on other classes, then based on the class names of the other classes it depends on, determine from multiple source code metadata the second target source code metadata whose file name is consistent with the class name of the other classes.
[0087] After determining the second target source code metadata, it is also possible to further determine whether the source code file corresponding to the second target source code metadata depends on other source code files. If it depends on other source code files, the above second target source code metadata determination process is repeated until the source code file corresponding to the determined target source code metadata does not depend on other source code files. Since the determination process is the same, the second target source code metadata determination process is taken as an example for exemplary description in this embodiment.
[0088] In this embodiment, by judging whether the first target source code metadata contains the class name of the parent class, parent interface, custom class and / or custom enumeration, it is determined whether the source code file corresponding to the first target source code metadata depends on other source code files.
[0089] If the first target source code metadata contains the class name of the parent class, parent interface, custom class and / or custom enumeration, it is determined that the source code file corresponding to the first target source code metadata depends on other source code files, and the other source code files it depends on are the source code files corresponding to the class name of the parent class, parent interface, custom class and / or custom enumeration contained in the first target source code metadata. If the first target source code metadata does not contain the class name of the parent class, parent interface, custom class and / or custom enumeration, it is determined that the source code file corresponding to the first target source code metadata does not depend on other source code files.
[0090] It is worth noting that in actual applications, the source code files corresponding to some custom classes may be open source, and the source code files corresponding to these custom classes will not be included in the source code files of the target program. Therefore, in the specific implementation process, if there is no source code metadata matching the class name of a certain custom class among the multiple source code metadata corresponding to the program source code of the target program, it is determined that the custom class is not used to determine the second program source code, that is, there is no need to obtain the program source code in the source code file of the custom class.
[0091] It is easy to understand that the determination process of the first target source code metadata and the second target source code metadata is similar, and both are based on the class name to determine the target source code metadata matching the class name from multiple source code metadata. In this embodiment, the determination process of the first target source code metadata is taken as an example for illustration.
[0092] Optionally, first target source code metadata having a similarity greater than a set threshold can be determined from multiple source code metadata based on the similarity between the class name of at least one class used by the method corresponding to the first program source code and the file name and the class name of the target class respectively contained in multiple source code metadata.
[0093] In the specific implementation process, Figure 3 As shown, multiple source code metadata can be respectively embedded (embedd ing) vector encoding processing to generate multiple first embedding vectors corresponding to multiple source code metadata (for example: first embedding vector 1 corresponding to source code metadata 1, first embedding vector 2 corresponding to source code metadata 2, ... first embedding vector m corresponding to source code metadata m), and the multiple first embedding vectors are stored in the vector database. When it is necessary to determine the first target source code metadata matching at least one class used by the method corresponding to the first program source code from the multiple source code metadata, the class name of the at least one class is subjected to the same embedding vector encoding processing to generate a second embedding vector corresponding to the class name of the at least one class. Afterwards, the similarity between the second embedding vector and the multiple first embedding vectors in the database is calculated (for example: cosine similarity, Euclidean distance, etc.), and the corresponding first target source code metadata whose similarity is greater than a set threshold is determined from the multiple source code metadata. Specifically, the corresponding target first embedding vector whose similarity is greater than the set threshold is determined from the multiple first embedding vectors, and the source code metadata corresponding to the target first embedding vector is used as the first target source code metadata.
[0094] It should be noted that Figure 3 The illustrated unit test code generation process only illustrates the process of determining the first target source code metadata. In fact, the process of determining the second target source code metadata is similar to it, and will not be repeated here.
[0095] Finally, the first target source code metadata and the second target source code metadata are used as target source code metadata that matches the class name of at least one class used by the method corresponding to the first program source code, and the program source code contained in the source code file corresponding to the target source code metadata is the second program source code.
[0096] In this scheme, the dependencies between source code files are taken into account by recursively determining the source code metadata, thereby ensuring the integrity of the determined target source code metadata, that is, ensuring the integrity of the second program source code determined based on the target source code metadata, thereby ensuring the correctness of the generated target unit test code.
[0097] The following will describe in detail the unit test code generation device of one or more embodiments of the present invention. Those skilled in the art will appreciate that these devices can be configured using commercially available hardware components through the steps taught in this solution.
[0098] Figure 5 A schematic diagram of the structure of a unit test code generation device provided by an embodiment of the present invention, such as Figure 5 As shown, the device includes: an acquisition module 11, a processing module 12, and a generation module 13.
[0099] The acquisition module 11 is used to acquire the program source code of the target program in response to a unit test code generation request triggered by a user, wherein the unit test code generation request is used to request generation of a target unit test code corresponding to a first program source code in the program source code.
[0100] The processing module 12 is used to input the program source code and the first program source code into a large language model so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code; determine source code metadata corresponding to the program source code, wherein the amount of data corresponding to the source code metadata is less than the amount of data corresponding to the program source code; and determine target source code metadata matching the class name of the at least one class from the source code metadata;
[0101] The generating module 13 is used to input the second program source code corresponding to the target source code metadata into the large language model, so that the large language model generates the target unit test code corresponding to the first program source code, and the program source code includes the second program source code.
[0102] Optionally, the acquisition module 11 is specifically used to acquire multiple source code files of the target program; and determine the program source code of the target program according to the program source codes respectively included in the multiple source code files.
[0103] Correspondingly, the processing module 12 is specifically used to determine multiple source code metadata corresponding to the multiple source code files as source code metadata corresponding to the program source code; determine target source code metadata that matches the class name of the at least one class from the multiple source code metadata, and the program source code contained in the source code file corresponding to the target source code metadata is the second program source code.
[0104] Optionally, the processing module 12 is further specifically used to parse the file name and file address of any target source code file among the multiple source code files, and the class name of the target class defined in the target source code file, wherein the target class includes at least one of a parent class, a parent interface, an inner class, a custom class, and a custom enumeration; and determine the source code metadata corresponding to the target source code file based on the file name, the file address, and the class name of the target class.
[0105] Optionally, the processing module 12 is further specifically used to determine first target source code metadata from the multiple source code metadata according to the class name of the at least one class and the file name and the class name of the target class respectively contained in the multiple source code metadata; if the first target source code metadata contains the class name of the parent class, parent interface, custom class and / or custom enumeration, then according to the class name of the parent class, parent interface, custom class and / or custom enumeration contained in the first target source code metadata and the file name and the class name of the target class respectively contained in the multiple source code metadata, determine second target source code metadata from the multiple source code metadata, the second target source code metadata being the source code metadata of the source code files corresponding to other classes on which the at least one class depends; and use the first target source code metadata and the second target source code metadata as the target source code metadata matching the class name of the at least one class.
[0106] The processing module 12 is further specifically used to determine, from the multiple source code metadata, the first target source code metadata whose similarity is greater than a set threshold based on the similarity between the class name of the at least one class and the file name and the class name of the target class respectively contained in the multiple source code metadata.
[0107] The processing module 12 is further specifically used to perform embedding vector encoding processing on the multiple source code metadata to generate multiple first embedding vectors corresponding to the multiple source code metadata; perform the embedding vector encoding processing on the class name of the at least one class to generate a second embedding vector corresponding to the class name of the at least one class; calculate the similarity between the second embedding vector and the multiple first embedding vectors; and determine the first target source code metadata whose similarity is greater than a set threshold from the multiple source code metadata.
[0108] Optionally, the acquisition module 11 is further specifically configured to acquire the target source code file corresponding to the target source code metadata according to the file address in the target source code metadata.
[0109] Figure 5 The device shown can execute the steps introduced in the aforementioned embodiments. For detailed execution process and technical effects, please refer to the description in the aforementioned embodiments, which will not be repeated here.
[0110] In one possible design, the above Figure 5 The structure of the unit test code generating device shown can be implemented as an electronic device, such as Figure 6 As shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. The memory 21 stores executable codes, and when the executable codes are executed by the processor 22, the processor 22 can at least implement the unit test code generation method provided in the above embodiments.
[0111] In addition, an embodiment of the present invention provides a non-temporary machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the unit test code generation method provided in the aforementioned embodiment.
[0112] An embodiment of the present invention provides a computer program product, including: a computer program, when the computer program is executed by a processor of an electronic device, the processor is caused to execute the unit test code generation method provided in the above embodiment.
[0113] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Those of ordinary skill in the art may understand and implement the present invention without creative effort.
[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on such an understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product, and the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating unit test code, characterized in that: include: In response to a unit test code generation request triggered by a user, obtaining a program source code of a target program, wherein the unit test code generation request is used to request generation of a target unit test code corresponding to a first program source code in the program source code; Inputting the program source code and the first program source code into a large language model, so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code; Determining source code metadata corresponding to the program source code, wherein the amount of data corresponding to the source code metadata is smaller than the amount of data corresponding to the program source code; Determining, from the source metadata, target source metadata that matches the class name of the at least one class; A second program source code corresponding to the target source code metadata is input into the large language model so that the large language model generates a target unit test code corresponding to the first program source code, wherein the program source code includes the second program source code.
2. The method according to claim 1, characterized in that The step of obtaining the program source code of the target program includes: Get multiple source code files of the target program; Determining the program source code of the target program according to the program source codes respectively included in the plurality of source code files; The step of determining source code metadata corresponding to the program source code and determining target source code metadata matching the class name of the at least one class from the source code metadata includes: Determining a plurality of source code metadata corresponding to the plurality of source code files as source code metadata corresponding to the program source code; Target source code metadata matching the class name of the at least one class is determined from the multiple source code metadata, and the program source code contained in the source code file corresponding to the target source code metadata is the second program source code.
3. The method according to claim 2, characterized in that The determining a plurality of source code metadata corresponding to the plurality of source code files comprises: For any target source code file among the multiple source code files, parse the file name and file address of the target source code file, and the class name of the target class defined in the target source code file, where the target class includes at least one of a parent class, a parent interface, an inner class, a custom class, and a custom enumeration; The source code metadata corresponding to the target source code file is determined according to the file name, the file address and the class name of the target class.
4. The method according to claim 3, characterized in that The step of determining target source metadata matching the class name of the at least one class from the plurality of source metadata comprises: Determining first target source metadata from the plurality of source metadata according to the class name of the at least one class and the file name and the class name of the target class respectively included in the plurality of source metadata; If the first target source code metadata includes the class name of a parent class, a parent interface, a user-defined class, and / or a user-defined enumeration, then according to the class name of the parent class, the parent interface, the user-defined class, and / or the user-defined enumeration included in the first target source code metadata, and the file name and the class name of the target class respectively included in the plurality of source code metadata, second target source code metadata is determined from the plurality of source code metadata, the second target source code metadata being the source code metadata of the source code files corresponding to other classes on which the at least one class depends; The first target source code metadata and the second target source code metadata are used as target source code metadata matching the class name of the at least one class.
5. The method according to claim 4, characterized in that The determining of first target source code metadata from the plurality of source code metadata according to the class name of the at least one class and the file name and the class name of the target class respectively included in the plurality of source code metadata comprises: According to the similarity between the class name of the at least one class and the file name and the class name of the target class respectively included in the multiple source code metadata, the first target source code metadata whose similarity is greater than a set threshold is determined from the multiple source code metadata.
6. The method according to claim 5, characterized in that The determining, from the plurality of source code metadata, first target source code metadata having a similarity greater than a set threshold value according to the class name of the at least one class and the similarity between the file name and the class name of the target class respectively included in the plurality of source code metadata, comprises: Performing embedding vector encoding processing on the plurality of source metadata to generate a plurality of first embedding vectors corresponding to the plurality of source metadata; Performing the embedding vector encoding process on the class name of the at least one class to generate a second embedding vector corresponding to the class name of the at least one class; Calculating similarities between the second embedding vector and the plurality of first embedding vectors; A first target source code metadata whose similarity is greater than a set threshold is determined from the multiple source code metadata.
7. The method according to claim 3, characterized in that The method further comprises: According to the file address in the target source code metadata, a target source code file corresponding to the target source code metadata is obtained.
8. A unit test code generation device, characterized in that: include: An acquisition module, configured to acquire a program source code of a target program in response to a unit test code generation request triggered by a user, wherein the unit test code generation request is used to request generation of a target unit test code corresponding to a first program source code in the program source code; a processing module, configured to input the program source code and the first program source code into a large language model, so that the large language model outputs a class name of at least one class used by a method corresponding to the first program source code; Determining source code metadata corresponding to the program source code, wherein the amount of data corresponding to the source code metadata is smaller than the amount of data corresponding to the program source code; determining target source code metadata matching the class name of the at least one class from the source code metadata; A generation module is used to input a second program source code corresponding to the target source code metadata into the large language model so that the large language model generates a target unit test code corresponding to the first program source code, and the program source code includes the second program source code.
9. An electronic device, characterized in that: include: A memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the unit test code generation method as described in any one of claims 1 to 7.
10. A non-transitory machine-readable storage medium, characterized in that: The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the unit test code generation method according to any one of claims 1 to 7.