Software testing method based on large language model, electronic equipment and medium
By using large language models and index databases in software testing, the similarity and redundancy of test cases are calculated, and the redundancy of test cases in traditional testing methods is solved, and the test efficiency is improved and the test accuracy is enhanced.
Patent Information
- Application Number
- CN202510064769.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
Smart Images

Figure CN119988217A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of software testing, and in particular to a software testing method, electronic device and medium based on a large language model. Background Art
[0002] With the continuous iteration and increasing complexity of software functions, traditional software testing methods are facing increasing challenges. In traditional testing, testers usually rely on a large set of test cases to verify each function of the software one by one. However, with the increase in R&D needs, the test case set continues to expand, which brings the following problems: First, the surge in the number of test cases has led to a large amount of redundancy and duplication. Different requirements may involve similar functions, resulting in the written test cases actually testing the same functional points, wasting manpower and time, and reducing testing efficiency. Second, the huge test case set has placed a heavy burden on system resources. Executing a large number of redundant test cases consumes more time and computing resources, prolongs the test cycle, and cannot meet the needs of rapid iteration development.
[0003] In order to solve these problems, artificial intelligence technology, especially large language models (LLM), has been gradually introduced into the field of software testing in recent years. However, the application of existing large language models in software testing still faces certain limitations and is difficult to fully adapt to complex and changing testing requirements. Therefore, how to optimize the screening of test cases and reduce duplication and redundancy by combining large language models has become a core problem that needs to be solved urgently in current technology. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a software testing method, electronic device and medium based on a large language model, which reduces the redundancy and repetition of test cases during software testing by combining the large language model technology.
[0005] To solve the above technical problems, an embodiment of the present invention provides a software testing method based on a large language model, including: inputting user R&D requirements and a test case set into a large language model, matching them in an index database, and outputting the functional coverage between each test case in the test case set and the R&D requirements; wherein the R&D requirements include at least one function; inputting the test cases and their functional coverage into a large language model, calculating the functional coverage similarity and redundancy between the test cases; notifying the user for review of the test cases whose functional coverage similarity is greater than a first preset threshold and whose redundancy is greater than a second preset threshold, and updating the test case set based on the user processing results.
[0006] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the software testing method based on the large language model as described above.
[0007] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the software testing method based on a large language model as described above.
[0008] In an embodiment of the present invention, firstly, a large language model is combined with an index database to accurately calculate the coverage of test cases for each function in the R&D requirements, identify test cases with the same functions, and determine which test cases involve the same or similar functional areas. Subsequently, the redundancy between these test cases with the same functions is calculated to evaluate their similarity in test content, steps, and expected results. High redundancy indicates that test cases may be repeated in actual execution, resulting in a waste of resources. Through this method, redundant test cases that need to be merged or deleted can be more accurately identified, effectively reducing the number of redundant test cases, avoiding repeated testing of the same functional points, and significantly improving test efficiency.
[0009] In addition, after updating the test case set based on the user processing results, it also includes: based on the characteristics of the test cases processed by the user, calculating the evaluation value of each test case through the large language model; inputting the software version update description into the second machine learning model and the large language model, and screening out the test cases that meet the requirements of the software version update description; wherein the software version update description has a mapping relationship with the required range of the evaluation value. On the basis of removing redundant cases, this method analyzes the version update description and uses the large language model to accurately match the key functions and high-priority test content required for the current version, effectively avoiding the waste of resources of low-priority or irrelevant test cases, ensuring that the test process focuses on the functional points directly related to the version update, not only improving the test efficiency, but also meeting the needs of rapid iterative development for efficient testing and accurate verification, thereby ensuring the quality and stability of software updates. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0011] Figure 1 is a flowchart of a software testing method provided by an embodiment of the present application;
[0012] Figure 2 This is a flowchart of calculating function coverage in a software testing method provided in an embodiment of the present application;
[0013] Figure 3 This is a flowchart of screening test cases that meet the version update description in the software testing method provided in an embodiment of the present application;
[0014] Figure 4 is an operation flow chart of a software testing method provided by an embodiment of the present application;
[0015] Figure 5 It is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0016] With the rapid iteration and increasing complexity of software functions, traditional software testing methods are facing increasing challenges. Testers usually rely on a large number of test cases to verify the functions of the software one by one. However, with the continuous expansion of R&D needs, the test case set is growing exponentially, which brings the following problems:
[0017] 1. Redundancy and duplication of test cases: Different R&D requirements may involve similar functional points, resulting in test cases repeatedly covering the same functional areas, wasting resources and reducing testing efficiency.
[0018] 2. Resource consumption and extended testing cycle: A large set of test cases increases execution costs and system burden, extends the testing cycle, and is difficult to adapt to rapid iterative development needs.
[0019] 3. The difficulty of test case optimization increases: Due to the lack of intelligent tools, the screening, optimization and dynamic update process of test cases rely on manual work, which makes it difficult to meet complex and changing testing requirements.
[0020] In recent years, the introduction of artificial intelligence technology (especially large language models, LLMs) has provided new ideas for solving these problems. However, when applied to test case screening, existing technologies have the limitation of being difficult to accurately adapt to complex test requirements. Therefore, how to combine large language models to achieve redundant identification, screening optimization and dynamic adjustment of test cases has become a core technical problem that needs to be solved urgently.
[0021] To make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings. However, it will be appreciated by those skilled in the art that in the embodiments of the present invention, many technical details are proposed in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can be implemented. The division of the following embodiments is for the convenience of description, and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined and referenced with each other without contradiction.
[0022] An embodiment of the present invention relates to a software testing method based on a large language model, which can be applied to electronic devices (such as servers, developer terminals, mobile terminals, and embedded devices, etc.). The method includes: inputting user R&D requirements and test case sets into a large language model, matching in an index database, and outputting the functional coverage between each test case in the test case set and the R&D requirements; wherein the R&D requirements include at least one function; inputting the test case and its functional coverage into a large language model, calculating the functional coverage similarity and redundancy between the test cases; notifying the user of the test case whose similarity of functional coverage is greater than a first preset threshold and whose redundancy is greater than a second preset threshold for review, and updating the test case set based on the user processing result. In an embodiment of the present invention, firstly, the coverage of the test case for each function in the R&D requirements is accurately calculated by combining a large language model with an index database, identifying the test cases with the same functions, and determining which test cases involve the same or similar functional areas. Subsequently, the redundancy between these test cases with the same functions is calculated, and the similarity in the test content, steps and expected results is evaluated. High redundancy indicates that the test case may be repeated in actual execution, resulting in a waste of resources. This method can more accurately identify redundant test cases that need to be merged or deleted, effectively reduce the number of redundant test cases, avoid repeated testing of the same functional points, and significantly improve test efficiency. Finally, combined with user review, the test case set is dynamically updated to further optimize the quality and efficiency of the test cases.
[0023] The implementation details of the software testing method based on a large language model according to an embodiment of the present invention are specifically described below. The following content is only provided for the convenience of understanding the implementation details and is not necessary for implementing the present solution.
[0024] like Figure 1 As shown, in step 110, an index database is constructed based on pre-stored R&D requirement documents and test case sets.
[0025] It should be noted that the data used to construct the index database in this embodiment include but are not limited to R&D requirement documents, test case sets and test result data generated by the last test. Among them, the R&D requirement documents provide function descriptions, requirement priorities, constraints and version iteration information, which are the basis for extracting function point keywords and contexts, and provide directional guidance for the functional matching of test cases. The test case set contains content such as case descriptions, test steps, function coverage information and execution conditions, which are used to establish the semantic association between test case keywords and contexts, and at the same time associate the test case keywords with R&D requirement function points. The software test result data includes test status, number of defects found, defect severity and version association information, which provides a reference basis for the optimization and supplement of the test case context, and can reflect the performance of the test case in actual execution and its dynamic adaptability with the function point. These data form a dynamic association between test case keywords and R&D requirement function points and test case keywords and context information through keyword extraction, vectorization processing and the establishment of multi-level mapping relationships, which not only improves the accuracy of test case screening, but also lays a data foundation for the quantitative calculation of function coverage.
[0026] In a possible embodiment, the specific steps of constructing the index database in step 110 include: first, extracting keywords and context information of the keywords (such as functional description, test steps, defect information, etc.) from pre-stored R&D requirement documents, test case sets and test results, and converting them into vector representations through natural language processing technology; then, using the first machine learning model to establish a semantic mapping relationship (first mapping relationship) between the R&D requirement keywords and the test case set keywords, ensuring the consistency of the functional points and the test objectives, as well as the association relationship (second mapping relationship) between the test case keywords and their contexts, so that the test case description is more closely associated with the actual functional points; finally, combining the defect description, severity and test status information in the test results, further optimizing the association between the test case keywords and the functional modules; finally, the keywords, context and original data are structured and stored, and a dynamically updateable index database is formed based on the first mapping relationship and the second mapping relationship. Through the above steps, this embodiment realizes the deep semantic association between R&D requirements, test cases and test results, and provides efficient and accurate data support for subsequent redundant processing and accurate screening of test cases based on large language models.
[0027] In actual implementation, the first machine learning model is a machine learning model that achieves regression effects, and can accurately establish semantic associations between R&D requirements and test cases, and between test case keyword tone contexts based on input keywords and context vectors. Specifically, the model can be a decision tree model, a neural network model, a support vector machine, or other models suitable for regression tasks.
[0028] In step 120, the device inputs the R&D requirements and test case set input by the user into the large language model, matches them in the index database, and outputs the functional coverage between each test case in the test case set and the R&D requirements.
[0029] It should be noted that, in this embodiment, the R&D requirements may consist of a single function point or a combination of multiple function requirements, which may involve specific operation modules, behavior descriptions, input and output conditions, priorities, etc. In order to more comprehensively characterize complex requirements, the function coverage is represented as vector data, where each dimension corresponds to a specific function point in the R&D requirements, and the vector value reflects the coverage of the test case for the relevant function point, such as whether it is covered or the relative proportion of coverage.
[0030] In a possible embodiment, the specific steps of calculating the functional coverage in step 120 are as follows: Figure 2 As shown, it includes: step 121, extracting keywords from user R&D requirements; step 122, matching corresponding test cases and context information in the index database based on the keywords of the R&D requirements; step 123, inputting the matching results into the large language model, and calculating the functional coverage between each test case in the test case set and the R&D requirements. For example, first, extracting keywords related to core functions from the R&D requirements input by the user through natural language processing technology, such as operation behavior, input and output conditions, priority information, etc., to ensure that the keywords can fully express the core content of the R&D requirements. Then, using the extracted keywords to perform semantic matching in the index database, retrieve the test cases associated with the keywords and their context information. The context information includes the description of the test case, test steps, expected results, and covered functional points and historical test result data (such as defect data, verification status). Subsequently, the matched test cases and their context information are input into the large language model together with the R&D requirements, and the matching degree between the test cases and the R&D requirements is quantitatively calculated through the semantic analysis capability of the large language model. Common methods such as cosine similarity or Jaccard coefficient are used to finally generate a functional coverage score.
[0031] The above functional coverage score can not only reflect the coverage of each test case for the functional points of R&D requirements, but also assist in subsequent test optimization and test case screening. This process improves the accuracy and comprehensiveness of functional matching by integrating context information, historical test data and functional point descriptions, laying a solid foundation for subsequent steps (such as redundant test case processing and test resource optimization).
[0032] In step 130, the device inputs the test case and its function coverage into the large language model, and calculates the function coverage similarity and redundancy between the test cases. Specifically, the device first generates a function coverage vector for each test case according to the function coverage calculated in step 120, wherein each dimension corresponds to a specific function point in the R&D requirements, and the value represents the coverage degree of the corresponding function point by the test case. The function coverage similarity between the test cases is preliminarily obtained by similarity calculation (such as cosine similarity or Jaccard coefficient) between the test cases, which is used to indicate the similarity degree of the function points covered between the test cases. Subsequently, the semantic level of the test case is deeply analyzed using the large language model, including comparing the semantic similarity of the test case description, test steps and expected results, focusing on evaluating whether there is functional duplication. Semantic analysis not only considers the duplication degree of the test steps and the similarity of the expected results, but also introduces the associated context information to further judge the redundancy of the test case. Combining the function coverage similarity and the semantic analysis results, the device obtains the redundancy score through weighted calculation, which provides a key basis for subsequent screening and processing.
[0033] In step 140, the device notifies the user of the test cases whose similarity of function coverage is greater than the first preset threshold and whose redundancy is greater than the second preset threshold for review, and updates the test case set based on the user's processing results. Specifically, the device first screens the test case pairs whose function coverage similarity is greater than the first preset threshold, and these test cases have significant functional similarity; then, in the screening results, further screens the test case pairs whose redundancy is greater than the second preset threshold, and marks them as redundant test cases. The system generates a redundancy report, including information such as function coverage, function coverage similarity, redundancy score, test step description and expected results of the test case pairs, and notifies the user for review. The user can process the redundant test cases according to the actual test requirements, such as deleting redundant test cases, merging similar test cases or retaining special test cases. The system records the user's operation decisions and reasons. These data are not only used to update the test case set, but also as a basis for feedback optimization of the large language model and screening rules to ensure that subsequent screening is more efficient and more accurate. Finally, the device dynamically updates the test case set according to the user's operation results, and dynamically updates the keywords, context and function point association relationship of the index database in combination with the characteristics of the updated test case set. This dynamic update mechanism can reflect the latest status of the test case set, while optimizing the semantic matching accuracy of the index database, providing precise support for subsequent test optimization.
[0034] In addition, in order to match test cases that are more in line with the current version update description, as an optional embodiment, after step 140, a large language model is also used to perform an operation of intelligently screening test cases (step 150), such as Figure 3, specifically including: step 151, based on the characteristics of the test cases processed by the user, calculating the evaluation value of each test case through the large language model; step 152, inputting the software version update description into the second machine learning model and the large language model, and screening out the test cases that meet the requirements of the software version update description to ensure that the selected test cases can cover the key requirements of the version update.
[0035] In the above embodiment, there is a clear mapping relationship between the software version update description and the evaluation value range of the test case. The version update description usually contains priority information, functional module descriptions and other test requirements, which determine the key direction of test case screening. High-priority version updates may require test cases with higher evaluation values. These test cases usually cover key functional points and can verify core logic or solve high-severity defects; while low-priority versions may focus more on the integrity testing of basic functions, allowing test cases with moderate evaluation values to be screened, thereby balancing the investment of testing resources.
[0036] In step 151, the evaluation value of the test case is calculated by the large language model based on multi-dimensional features, including but not limited to: functional coverage (i.e., the breadth and depth of the test case covering the functional points of R&D requirements), effectiveness (the ability of the test case to find problems or verify functions), execution version information (the adaptability and execution effect of the test case in the historical version), and defect severity (the degree of impact of the defects found in the historical execution of the test case on the system function). The large language model comprehensively analyzes these features and generates a dynamic evaluation value reflecting the importance of the test case through weighted or dynamic learning, thereby more accurately supporting the screening of test cases.
[0037] In step 152 of the above embodiment, the screening of test cases is completed by combining the machine learning model and the large language model, specifically including the following steps: the priority of the software version is input into the second machine learning model, and the model outputs the evaluation value range corresponding to the current version according to the mapping relationship between the pre-trained priority and the evaluation value range. The priority in the version update description determines the criticality of this version test and the allocation direction of test resources. For example, a high-priority version may be mapped to a range of evaluation values of 85 and above, while a medium- and low-priority version may cover an evaluation value range between 60 and 80. Subsequently, the version update description and its corresponding evaluation value range are input into the large language model, and its semantic analysis and context matching capabilities are used to screen qualified test cases from the index database. This process not only relies on the keyword analysis of the version update description, but also comprehensively considers the historical performance, functional coverage, and defect relevance of the test case, so as to screen out the test case that best meets the version requirements and maximize the utilization of test resources.
[0038] In order to continuously optimize the screening accuracy and adapt to the dynamic changes of the software version, after step 152, the parameters of the machine learning model and the large language model need to be dynamically updated, specifically including: first, the software version is tested based on the screened test cases to generate test result data; then, the test results are input into the large language model, and the evaluation value of the test case is updated to ensure that the evaluation value can reflect the latest test results and defect discovery; finally, the test results are simultaneously input into the second machine learning model to update the mapping relationship between the version priority and the evaluation value. This parameter update mechanism enables the large language model and the machine learning model to be continuously optimized and better adapt to the dynamic needs in the software version iteration.
[0039] In addition, the updated second machine learning model and large language model can continue to be applied to the new version of the software test case screening process. The specific process includes: re-acquiring the version update description input by the user; recalculating the evaluation value of the screened test case based on the updated large language model; inputting the priority of the software version into the updated second machine learning model, and re-outputting the evaluation value range corresponding to the priority; inputting the version update description, evaluation value range and test results into the large language model, and re-screening the test cases that meet the version requirements from the index database. Through this continuous iterative optimization process, the system can dynamically adapt to software update requirements and provide more efficient and accurate support for testing.
[0040] In addition, in a possible implementation, after step 140, the device will also input the test case set reviewed and processed by the user into the large language model to dynamically update the second preset threshold. Specifically, the device analyzes user preferences and actual needs based on the user's operation data when processing redundant test cases (such as decisions to delete, merge or retain) and the functional coverage similarity and redundancy information of related test cases. Based on these data, the device adjusts the setting logic of the second preset threshold to enable it to more accurately identify and filter redundant test cases. At the same time, the device stores the optimized threshold and applies it to subsequent screening, so that the screening process is more in line with actual testing needs, further improving the screening efficiency and the rationality of the test case set.
[0041] In an embodiment of the present invention, firstly, a large language model is combined with an index database to accurately calculate the coverage of test cases for each function in the R&D requirements, identify test cases with the same functions, and determine which test cases involve the same or similar functional areas. Subsequently, the redundancy between these test cases with the same functions is calculated to evaluate their similarity in test content, steps, and expected results. High redundancy indicates that test cases may be repeated in actual execution, resulting in a waste of resources. Through this method, redundant test cases that need to be merged or deleted can be more accurately identified, effectively reducing the number of redundant test cases, avoiding repeated testing of the same functional points, and significantly improving test efficiency. Finally, combined with user review, the test case set is dynamically updated to further optimize the quality and efficiency of the test cases.
[0042] In another possible embodiment, to adapt to different test demand scenarios, the intelligent screening test case operation from step 151 to step 152 can be run as an independent process without relying on the redundant test case intelligent identification method from step 130 to step 140. This design is a flexible choice of implementation method, and can be flexibly combined according to actual needs during specific implementation.
[0043] like Figure 4 As shown in Figure 1, the architecture of the intelligent language model consists of two main modules: data preparation process and test case screening process. The data preparation process shows the complete process from data collection to vector database construction, including R&D requirement documents, test case sets, and test result data extraction, vectorization, and index relationship establishment; the test case screening process shows the specific process from retrieving data from the index database to processing the large language model or machine learning model.
[0044] exist Figure 4On the left side of the figure, it shows in detail how to use the large language model to perform intelligent screening of test cases. When the user is ready to release a new version of the software, the redundant identification stage can be skipped and the software version update description and the test case set can be directly input into the intelligent language model. The intelligent language model operates in coordination with the large language model and the machine learning model therein, and performs the operations of step 151 and step 152 in sequence. In step 151, based on the input test case features (such as functional coverage, effectiveness, execution version information, and defect severity), the large language model calculates the dynamic evaluation value of each test case; in step 152, combined with the input of the software version update description, the intelligent language model uses the second machine learning model to generate a mapping between priority and evaluation value range, and uses the large language model to filter out the test case set that is most suitable for the current software version from the index database. This screening process ensures that the test case can accurately match the functional requirements and priority requirements of the current version. Based on the screened test case set, the user can directly perform automated testing or scenario testing, thereby quickly generating test result data for the current software version and completing the test process. In addition, the generated test result data can also be used as an important input for subsequent optimization of the intelligent language model, providing support for the model to dynamically adjust the screening logic and evaluation rules.
[0045] At the same time, Figure 4 On the right side, when the user needs to perform intelligent identification of redundant test cases, the R&D requirements and test case sets can be input into the intelligent language model, triggering the large language model and the machine learning model to perform operations from step 130 to step 140. Specifically, in step 130, the intelligent language model calculates the functional coverage similarity and redundancy between test cases; in step 140, the test cases whose redundancy exceeds the preset threshold are screened out, and these test cases are automatically marked as redundant test cases, and the user is notified to review or automatically perform processing (such as deletion, merging or retention). Users can further adjust the direction of software development according to actual needs, update R&D requirements and test case sets, to ensure the accuracy and efficiency of test cases.
[0046] Through the above embodiments, the intelligent language model achieves the decoupling of two core functions: on the one hand, it supports the rapid screening of the test case set that best meets the current version requirements and directly puts it into the test process; on the other hand, it supports the intelligent identification and optimization of redundant test cases, providing strong support for R&D demand adjustment and test case optimization. This flexible architecture design not only improves the efficiency and accuracy of software testing, but also can dynamically adapt to different test demand scenarios, providing comprehensive support for the rapid iteration of software development processes.
[0047] The steps of the above method are divided only for the purpose of clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.
[0048] In addition, the examples mentioned in the above embodiments can be freely combined, and any combination can be understood as an embodiment. The "embodiment" or "example" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It can be understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0049] Another embodiment of the present invention relates to an electronic device, such as Figure 5 As shown, it includes at least one processor 501; and a memory 502 that is communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the software testing method based on the large language model as described above.
[0050] Among them, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium via an antenna, and further, the antenna also receives data and transmits the data to the processor.
[0051] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.
[0052] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program, which implements the above method embodiment when executed by a processor.
[0053] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0054] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present invention, and that in actual applications, various changes may be made in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A software testing method based on a large language model, characterized in that: include: Inputting the user's R&D requirements and the test case set into the large language model, matching them in the index database, and outputting the functional coverage between each test case in the test case set and the R&D requirements; wherein the R&D requirements include at least one function; Inputting the test cases and their functional coverage into a large language model, and calculating the functional coverage similarity and redundancy between the test cases; The test cases whose similarity of function coverage is greater than a first preset threshold and whose redundancy is greater than a second preset threshold are notified to the user for review, and the test case set is updated based on the user processing results.
2. The software testing method based on a large language model according to claim 1, characterized in that: Before inputting the user R&D requirements and the test case set into the large language model, the method further includes: Extract keywords and contexts of the keywords from pre-stored R&D requirement documents and test case sets, and convert them into vector representations; Using a first machine learning model, establish a first mapping relationship between the R&D requirement keywords and the test case set keywords, and a second mapping relationship between the test case set keywords and context; The index database is constructed based on the first mapping relationship and the second mapping relationship.
3. The software testing method based on a large language model according to claim 2, characterized in that: The step of inputting the user R&D requirements and the test case set into the large language model, matching them in the index database, and outputting the functional coverage between each test case in the test case set and the R&D requirements specifically includes: Extract keywords from user R&D needs; Based on the keywords of the R&D requirements, corresponding test cases and context information are matched in the index database; The matching results are input into a large language model to calculate the functional coverage between each test case in the test case set and the R&D requirements.
4. The software testing method based on a large language model according to any one of claims 1 to 3, characterized in that: After the test case set is updated based on the user processing result, the method further includes: The updated test case set is input into the large language model, and the second preset threshold is updated.
5. The software testing method based on a large language model according to any one of claims 1 to 3, characterized in that: After updating the test case set based on the user processing result, the method further includes: Based on the features of the test cases processed by the user, calculating the evaluation value of each test case through the large language model; The software version update description is input into the second machine learning model and the large language model to filter out test cases that meet the requirements of the software version update description; wherein the software version update description has a mapping relationship with the required range of the evaluation value.
6. The software testing method based on a large language model according to claim 5, characterized in that: The step of inputting the software version update description into the second machine learning model and the large language model to filter out test cases that meet the requirements of the software version update description specifically includes: Inputting the priority of the software version into a second machine learning model, and outputting a range of evaluation values corresponding to the priority; wherein the version update description includes the priority; The version update description and the evaluation value range are input into the large language model, and test cases that meet the version update description and the evaluation value range are screened from the index database.
7. The software testing method based on a large language model according to claim 6, characterized in that: After screening the test cases that meet the version update description and the evaluation value range from the index database, the method further includes: Performing a test on the software version based on the screened test cases to generate a test result; Inputting the test results into the large language model, and updating the evaluation values of the screened test cases; The test result is input into the second machine learning model to update the mapping relationship between the version priority and the evaluation value.
8. The software testing method based on a large language model according to claim 7, characterized in that: After the updating of the mapping relationship between the version priority and the evaluation value, the method further includes: Re-obtain the version update description entered by the user; Recalculating the evaluation values of the screened test cases based on the updated large language model; Inputting the priority of the software version into the updated second machine learning model, and re-outputting the evaluation value range corresponding to the priority; The version update description, the evaluation value range and the test result are input into the large language model, and the test cases that meet the version update description and the evaluation value range are re-screened from the index database.
9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the software testing method based on a large language model as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the software testing method based on a large language model according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Test case generation method and device, equipment and computer program product
CN121070344A
Test case batch generation and iterative self-repairing method based on large model agent
CN122195860A