Non-functional test method and device of software system and electronic equipment
By using pre-trained large language models to parse the requirements document and obtain non-functional testing rules from the knowledge base, and generate non-functional testing products, it solves the problems of low efficiency and low quality of non-functional testing of software systems in the prior art, and achieves efficient and accurate non-functional testing.
Patent Information
- Application Number
- CN202510533082.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art has problems of low efficiency, low quality and experience dependence in non-functional testing of software systems. Automation methods fail to fully consider the special needs of non-functional testing and may cause hallucinations.
Use pre-trained large language models to parse and improve requirements documents, extract non-functional testing key points and requirements details, obtain target non-functional testing rules from the pre-built knowledge base, generate non-functional testing products, and ensure the test quality by evaluating the integrity, correctness and executability of the products.
It significantly improves the efficiency and quality of non-functional testing, reduces dependence on manual experience, and ensures the comprehensiveness and accuracy of the testing process.
Smart Images

Figure CN120448266A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of software testing, and in particular to a non-functional testing method for a software system, a non-functional testing device for a software system, a computer-readable storage medium, and an electronic device. Background Art
[0002] Non-functional testing of software systems requires work planning, manual editing, verification and review, release of test assets, and test execution, all within the requirements analysis, test planning, test case development, test execution, results analysis, and test summary stages. Traditionally, these tasks are performed by personnel after understanding the requirements and test specifications. This results in a time-consuming process, a dependence on individual experience for the quality of the results, and low efficiency.
[0003] With the development of artificial intelligence (AI), particularly the application of large models in deep learning, existing technologies are leveraging AI to improve software testing processes. However, some existing solutions, such as rapid implementation of result judgment based on image recognition and transfer learning, or intelligent use case generation methods, focus on automation or test case generation in specific areas, failing to fully consider the unique needs of non-functional testing and failing to address the potential for hallucinations caused by large models in testing, where the content generated by the model may not match the actual situation.
[0004] In summary, although existing technologies have introduced some automation methods when dealing with non-functional testing of software systems, they still cannot effectively overcome key problems such as time-consuming, reliance on personal experience, low efficiency and poor quality of test products. Summary of the Invention
[0005] The main purpose of this application is to provide a non-functional testing method for a software system, a non-functional testing device for a software system, a computer-readable storage medium and an electronic device, so as to at least solve the problems of limitations of the non-functional testing methods of software systems in the prior art, including low efficiency, low quality and dependence on experience.
[0006] To achieve the above-mentioned objectives, according to one aspect of the present application, a non-functional testing method for a software system is provided, the non-functional testing method including multiple types of non-functional testing, each type of non-functional testing including corresponding non-functional testing key points and requirement details, including: obtaining an improvement requirement document for the software to be tested, and parsing the improvement requirement document using a pre-trained large language model to extract non-functional testing key points and requirement details, wherein the improvement requirement document refers to a document including areas where the functions and performance of the software to be tested need to be improved; obtaining target non-functional testing rules that match the non-functional testing key points and the requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes multiple non-functional testing rules; generating non-functional test artifacts using the pre-trained large language model according to the target non-functional testing rules, wherein the non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports; and performing non-functional testing on the software to be tested using the non-functional test artifacts if the non-functional test artifacts meet preset standards, wherein the preset standards are formulated based on at least one of the following: integrity, correctness, and executability of the non-functional test artifacts.
[0007] Optionally, the non-functional test rules are in vector form, and target non-functional test rules that match the non-functional test points and the requirement details are obtained from a pre-built knowledge base, and the pre-built knowledge base includes multiple non-functional test rules, including: converting the non-functional test points and the requirement details into vector representations to obtain non-functional test point vectors and requirement detail vectors; calculating the similarity between the non-functional test point vectors and the requirement detail vectors and each of the non-functional test rules in the pre-built knowledge base to obtain multiple similarity results; based on the multiple similarity results, using a top-k recall algorithm to sort all the non-functional test rules to obtain the top k non-functional test rules with the highest similarity to the non-functional test point vectors and the requirement detail vectors, and determining the top k non-functional test rules as the target non-functional test rules, where k≥1.
[0008] Optionally, before obtaining target non-functional test rules that match the non-functional test points and the requirement details from a pre-built knowledge base, and the pre-built knowledge base includes multiple non-functional test rules, the method further includes: obtaining historical non-functional test cases and first expert experience data related to the non-functional test, the historical non-functional test cases including historical non-functional test cases, historical non-functional test plans, historical non-functional test report results and historical non-functional test results; generating multiple non-functional test rules based on the historical non-functional test cases and the first expert experience data using the pre-trained large language model; using a word segmenter to split each of the non-functional test rules into multiple word segmentation units; inputting the multiple word segmentation units into a pre-built Transformer model to generate multiple word vectors, and updating each of the word vectors based on a self-attention mechanism to obtain multiple initial vectors, the initial vectors including relevant information between each of the word segmentation units about the text information of the non-functional test rules; performing a pooling operation on the multiple initial vectors to generate multiple non-functional test rule vectors, and obtaining the pre-built knowledge base composed of multiple non-functional test rule vectors.
[0009] Optionally, according to the target non-functional test rules, non-functional test products are generated using the pre-trained large language model, and the non-functional test products include non-functional test cases, non-functional test plans and non-functional test reports, including: generating prompt words according to the target non-functional test rules and second expert experience data, and the prompt words include: generating test cases from the following five dimensions, the first dimension is to perform a single-interface performance benchmark test on each test interface information, the second dimension is to perform a single-interface load test on each of the test interface information, the third dimension is to perform a single-interface green light test on each of the test interface information, the fourth dimension is to perform a mixed transaction capacity test on all interfaces, and the fifth dimension is to perform a mixed transaction fatigue test on all the interfaces; the prompt words are input into the pre-trained large language model, and combined with the pre-built knowledge base to generate the non-functional test products.
[0010] Optionally, an improvement requirement document of the software to be tested is obtained, and a pre-trained large language model is used to parse the improvement requirement document to extract non-functional test points and requirement details, including: obtaining the improvement requirement document of the software to be tested; determining whether the improvement requirement document includes non-functional requirements; if the improvement requirement document does not include the non-functional requirements, outputting a prompt message and ending the process; if the improvement requirement document includes the non-functional requirements, inputting the improvement requirement document into the pre-trained large language model, and generating the non-functional test points and the requirement details according to preset prompt words.
[0011] Optionally, after generating a non-functional test product using the pre-trained large language model according to the target non-functional test rules, the method further includes: if the non-functional test product does not meet the preset standard, adjusting the prompt words input into the pre-trained large language model and relevant parameters of the pre-trained large language model, and regenerating the non-functional test product, wherein the relevant parameters include learning rate, regularization parameter, context length and generation length.
[0012] Optionally, after generating non-functional test products using the pre-trained large language model according to the target non-functional test rules, the method further includes: obtaining offline evaluation results and online user feedback on the non-functional test products; identifying the non-functional test products with a recall rate lower than a preset threshold based on the offline evaluation results and the online user feedback; converting the positive examples of the non-functional test products with a recall rate lower than the preset threshold into enhanced non-functional test rules to enhance the ability of the pre-trained large language model to generate accurate non-functional test products, and converting the negative examples of the non-functional test products with a recall rate lower than the preset threshold into error-correcting non-functional test rules to guide the pre-trained large language model to identify potential error generation patterns and take measures to prevent the potential errors from occurring again.
[0013] According to another aspect of the present application, a non-functional testing apparatus for a software system is provided, comprising: a parsing unit for obtaining an improvement requirement document for software to be tested, and parsing the improvement requirement document using a pre-trained large language model to extract non-functional testing key points and requirement details, wherein the improvement requirement document refers to a document including areas where the functions and performance of the software to be tested need to be improved; a first acquisition unit for obtaining target non-functional testing rules that match the non-functional testing key points and the requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes a plurality of non-functional testing rules; a first generation unit for generating non-functional test artifacts using the pre-trained large language model according to the target non-functional testing rules, wherein the non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports; and a non-functional testing unit for performing non-functional testing on the software to be tested using the non-functional test artifacts if the non-functional test artifacts meet preset standards, wherein the preset standards are formulated based on at least one of the following: integrity, correctness, and executability of the non-functional test artifacts.
[0014] According to another aspect of the present application, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute any one of the non-functional testing methods for the software system.
[0015] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a non-functional testing method for executing any one of the software systems described.
[0016] Apply the technical solution of the present application to obtain an improvement requirement document of the software to be tested, and use a pre-trained large language model to parse the improvement requirement document to extract non-functional test key points and requirement details. The improvement requirement document refers to a document that includes the functions and performance that need to be improved in the software to be tested; obtain target non-functional test rules that match the non-functional test key points and requirement details from a pre-built knowledge base, and the pre-built knowledge base includes a variety of non-functional test rules; according to the target non-functional test rules, use the pre-trained large language model to generate non-functional test artifacts, and the non-functional test artifacts include non-functional test cases, non-functional test plans and non-functional test reports; if the non-functional test artifacts meet the preset standards, use the non-functional test artifacts to perform non-functional testing on the software to be tested, wherein the preset standards are formulated based on at least one of the following: the completeness, correctness and executability of the non-functional test artifacts. This solution uses a pre-trained large language model to parse and transform requirement documents, extract the key points and requirement details of non-functional testing, and then obtain relevant non-functional testing rules from a pre-built knowledge base to further guide the pre-trained large language model to generate non-functional test artifacts, significantly improving the efficiency of non-functional testing and reducing dependence on manual experience. By judging whether the non-functional test artifacts meet the preset standards, the quality of non-functional testing is improved, thereby solving the limitations of the existing methods of non-functional testing of software systems, including low efficiency, low quality and experience dependence. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:
[0018] Figure 1 A hardware structure block diagram of a mobile terminal for executing a non-functional testing method of a software system provided in an embodiment of the present application is shown;
[0019] Figure 2 A flowchart of a non-functional testing method for a software system provided in accordance with an embodiment of the present application is shown;
[0020] Figure 3A flowchart of a specific non-functional testing method for a software system provided according to an embodiment of the present application is shown;
[0021] Figure 4 A flowchart of generating non-functional test artifacts based on a large language model in a non-functional testing method for a software system provided in an embodiment of the present application is shown;
[0022] Figure 5 The figure shows a structural block diagram of a non-functional testing device for a software system provided according to an embodiment of the present application.
[0023] The above drawings include the following reference numerals:
[0024] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. DETAILED DESCRIPTION
[0025] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] As introduced in the background technology, although the existing technology has introduced some automation means when dealing with non-functional testing of software systems, it is still unable to effectively overcome some limitations of non-functional testing. In order to solve the problems of limitations of the methods of non-functional testing of software systems in the existing technology, including low efficiency, low quality and experience dependence, the embodiments of the present application provide a non-functional testing method for a software system, a non-functional testing device for a software system, a computer-readable storage medium and an electronic device.
[0029] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0030] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure diagram of a mobile terminal of a non-functional testing method for a software system according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0031] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the device information display method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0032] In this embodiment, a non-functional testing method for a software system running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] Figure 2 FIG. 1 is a flow chart of a non-functional testing method for a software system according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:
[0034] Step S201: Obtain an improvement requirement document for the software to be tested, and use a pre-trained large language model to parse the improvement requirement document to extract non-functional test points and requirement details. The improvement requirement document refers to a document that includes the areas where the functions and performance of the software to be tested need to be improved.
[0035] Specifically, an improvement requirements document is a document that documents needed improvements or additions to software functionality and performance during the software maintenance, upgrade, or development phase. This type of document typically includes feedback from software engineers, product managers, or testers on existing software features, or plans and requirements for future functionality and performance improvements. The document details which parts of the software need to be enhanced, how performance will be optimized, and any non-functional features that may be involved, such as security, stability, response time, resource consumption, and compatibility.
[0036] Pre-trained large language models are trained on extensive data and possess powerful language understanding and generation capabilities. In non-functional testing scenarios, these pre-trained large language models are used to parse and improve requirements documents, aiming to automatically identify and understand descriptions of non-functional features within the documents, including specific performance metrics, security requirements, and stability standards. Leveraging the natural language processing capabilities of the large language model, unstructured document content is converted into structured data, facilitating subsequent non-functional testing rule matching and artifact generation.
[0037] Non-functional testing key points refer to test items or metrics that require special attention during non-functional testing, such as response time and throughput in performance testing, data leakage risk points in security testing, and system crash points in stability testing. The requirements specification provides detailed requirements or descriptions of these test key points, such as the specific response time indicators to be achieved in performance testing and the types of sensitive data to be checked in security testing.
[0038] After parsing the improvement requirements document, the pre-trained large language model extracts key points (non-functional test points) and specific requirements (requirement details) related to non-functional features in the improvement requirements document. This process can be thought of as converting the improvement requirements of the software under test into test specifications, guiding the subsequent non-functional test design and execution. For example, if the improvement requirements document mentions a performance bottleneck, the pre-trained large language model will identify this and analyze the various indicators and requirements related to performance testing in detail, ensuring that non-functional testing can comprehensively evaluate the performance improvement effects of the software under test.
[0039] Software non-functional testing involves a variety of types, including performance testing, stress testing, stability testing, compatibility testing, and security testing. Each type of non-functional testing has its own specific key points and detailed requirements. For example, performance testing focuses on response time and concurrent processing capabilities, while security testing focuses on data protection and attack resistance. Pre-trained large language models must be able to understand and distinguish these different types of non-functional testing requirements, accurately extracting the key points and detailed requirements for each test type, and providing precise guidance for the subsequent generation of non-functional test artifacts.
[0040] In summary, using a pre-trained large language model to parse and improve requirements documents and extract key non-functional test points and requirement details can significantly improve the efficiency, comprehensiveness, and product quality of non-functional testing. This process fully utilizes the natural language understanding and generation capabilities of the large language model, overcoming the limitations of traditional non-functional testing that relies on manual understanding and providing strong data support for automated and intelligent non-functional testing processes.
[0041] Step S202: acquiring target non-functional test rules that match the non-functional test key points and the requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes a plurality of non-functional test rules;
[0042] Specifically, a knowledge base is a database system that stores and manages a large amount of structured or semi-structured information. The purpose of building a knowledge base is to provide professional knowledge and data support in the field. In the embodiment of the present application, the knowledge base is specifically used in the field of software non-functional testing and includes a variety of non-functional testing rules. After the pre-trained large language model extracts the non-functional testing key points and requirement details from the improvement requirements document, the information in the knowledge base is used to match and obtain the most suitable non-functional testing rules, namely the target non-functional testing rules.
[0043] Through step S202, based on the rich information of the improved requirement document and knowledge base, non-functional testing rules that match the non-functional testing points and requirement details can be intelligently and accurately obtained from the pre-built knowledge base, thereby improving the generation quality and pertinence of non-functional testing products and comprehensively improving the efficiency and quality of software non-functional testing.
[0044] Step S203: generating non-functional test artifacts using the pre-trained large language model according to the target non-functional test rules, wherein the non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports;
[0045] Specifically, the pre-trained large language model has the ability to process and understand large amounts of text data. By self-learning from massive corpora, it can grasp the structure, grammar, and semantics of a language, as well as knowledge from various professional fields. In this embodiment, this pre-trained large language model is used to generate non-functional test artifacts. Based on its powerful generation capabilities, text content can be created based on input instructions or rules.
[0046] Targeted non-functional testing rules are a set of non-functional testing rules selected from a pre-built knowledge base that best matches the non-functional testing key points and requirements of the software to be tested. These non-functional testing rules include specific test types (such as performance testing, security testing, and compatibility testing), test steps, common testing methods, and expert experience accumulated from previous testing. The selection of target non-functional testing rules is based on the analysis results of the improved requirements document, ensuring that the generated non-functional test artifacts are directly targeted at the non-functional testing requirements of the software.
[0047] The pre-trained large language model generates non-functional test cases based on the target non-functional test rules. These non-functional test cases are specific operational instructions or data inputs designed to test the software's performance under specific non-functional metrics. For example, for performance testing, non-functional test cases may include simulating requests under different concurrent user numbers to evaluate the software's response time and processing capabilities. For security testing, non-functional test cases may involve attempts to inject malicious data or access sensitive resources to test the software's security mechanisms.
[0048] The pre-trained large language model can also generate detailed non-functional test plans. A non-functional test plan is a plan for the entire testing process, including test objectives, test environment preparation, test steps, expected results, and test resource requirements. The non-functional test plan generated by the pre-trained large language model ensures systematic and comprehensive testing, providing guidance on how to execute tests and evaluate test results.
[0049] The pre-trained large language model can generate a non-functional test report. This report is essentially a predictive report. Based on the details of the non-functional test cases and non-functional test plan, this report predefines information such as expected test results, potential risk points, an overview of the test strategy, and the test environment configuration. This predictive report helps the test team understand the test plan in advance and be fully prepared for test execution.
[0050] Through step S203, the pre-trained large language model can generate a series of non-functional test artifacts, including non-functional test cases, non-functional test plans, and non-functional test reports. This process not only significantly improves the efficiency and automation of non-functional testing, but also ensures the quality and professionalism of test artifacts through the large model's generation capabilities. Furthermore, because artifact generation is based on specific target non-functional test rules, the testing process is more targeted, better meeting the needs of software non-functional testing and enhancing the intelligence and standardization of the testing process.
[0051] Step S204: If the non-functional test artifact meets the preset standard, the non-functional test artifact is used to perform non-functional testing on the software to be tested, wherein the preset standard is formulated based on at least one of the following: integrity, correctness and executability of the non-functional test artifact.
[0052] Specifically, during the software non-functional testing process, the non-functional test artifacts generated by the pre-trained large language model, including non-functional test cases, non-functional test plans, and non-functional test reports, undergo a series of evaluations to ensure they meet pre-set standards. These pre-set standards are based on the completeness, correctness, and enforceability of the non-functional test artifacts, ensuring the effectiveness and reliability of the testing process.
[0053] The completeness of non-functional test artifacts refers to whether the non-functional test artifacts cover all necessary test points and requirements, and whether the test plan thoroughly describes each test step and expected results. When evaluating the completeness of non-functional test artifacts, check whether the non-functional test artifacts contain all relevant information such as test types, test cases, and test environment configurations identified based on the improved requirements document and target test rules. For example, performance test artifacts should list in detail all performance indicators that need to be tested, as well as how to design test cases to cover these indicators. The completeness of non-functional test artifacts that meet preset standards can ensure that the test process does not miss important test points, thereby improving the comprehensiveness and effectiveness of the test.
[0054] The accuracy of non-functional test artifacts focuses on whether the information in the artifacts is accurate, whether the test cases truly reflect the key points of non-functional testing, and whether the steps in the test plan are correct. This assessment involves reviewing the artifacts to ensure that the test cases and plans are free of logical errors, incorrect test parameters or test data, and that the test environment is properly configured. Meeting pre-set standards for the accuracy of non-functional test artifacts prevents ineffective or misleading testing due to artifact errors, ensuring the accuracy and reliability of test results.
[0055] The executability of non-functional test artifacts considers their operability and feasibility in the actual test environment. This ensures that the test plan can be executed, that the test cases can be correctly run by the test tool or framework, and that the expected results analysis in the test report can be verified. It is necessary to check whether each step in the test plan is supported by the actual environment, whether the test cases can be successfully executed by the test tool, and whether the analysis in the test report is based on verifiable test data and results. Non-functional test artifact executability that meets preset standards ensures a smooth test process and avoids wasting time and resources due to artifact non-executability.
[0056] Once a non-functional test artifact has passed the pre-defined criteria for completeness, correctness, and executable, it can be used in actual non-functional testing. The generated non-functional test plan is used to prepare the test environment, execute non-functional test cases, and generate a complete non-functional test report based on the actual test results. During this process, the completeness of the non-functional test artifact ensures that all test points are covered, its correctness ensures accurate testing, and its executable ensures the smooth implementation of the test plan, all contributing to the successful execution of the non-functional testing process.
[0057] Through this embodiment, a pre-trained large language model is used to parse and transform the requirement document, extract the key points and requirement details of non-functional testing, and then obtain relevant non-functional testing rules from a pre-built knowledge base to further guide the pre-trained large language model to generate non-functional test artifacts, which significantly improves the efficiency of non-functional testing and reduces dependence on manual experience. By judging whether the non-functional test artifacts meet the preset standards, the quality of non-functional testing is improved, thereby solving the limitations of the existing methods of non-functional testing of software systems, including low efficiency, low quality and experience dependence.
[0058] During the specific implementation process, the above-mentioned non-functional test rules are in vector form, and target non-functional test rules that match the above-mentioned non-functional test points and the above-mentioned requirement details are obtained from a pre-built knowledge base. The above-mentioned pre-built knowledge base includes a variety of non-functional test rules, including: converting the above-mentioned non-functional test points and the above-mentioned requirement details into vector representations to obtain non-functional test point vectors and requirement detail vectors; calculating the similarity between the above-mentioned non-functional test point vectors and the above-mentioned requirement detail vectors and each of the above-mentioned non-functional test rules in the above-mentioned pre-built knowledge base respectively, and obtaining multiple similarity results; based on the multiple similarity results, using the top-k recall algorithm to sort all the above-mentioned non-functional test rules, and obtain the top k non-functional test rules with the highest similarity to the above-mentioned non-functional test point vectors and the above-mentioned requirement detail vectors, and determining the top k non-functional test rules as the above-mentioned target non-functional test rules, where k≥1.
[0059] Specifically, a pre-trained language model such as Transformer can be used to convert the non-functional test key points and requirement details into vector representations to obtain non-functional test key point vectors and requirement detail vectors. Through vectorization, non-functional test key points and requirement details can be understood and processed by the computer system, providing a basis for subsequent similarity calculations. Similarity calculations can measure the degree of similarity between two vectors and can be implemented through a variety of distance or similarity measurement methods, such as cosine similarity and Euclidean distance. Similarity calculations are performed on the non-functional test key point vectors and requirement detail vectors respectively with each non-functional test rule vector in the pre-built knowledge base. For example, using cosine similarity, the cosine value of the angle between the two vectors is calculated. The closer the value is to 1, the more similar the two are. The results of the similarity calculation can help identify which non-functional test rules are most compatible with the current non-functional test key point vectors and requirement details, thereby optimizing the selection of non-functional test rules.
[0060] After obtaining the similarity results, the top-k recall algorithm is used to sort all non-functional test rules. The top-k recall algorithm is an algorithm used in the field of information retrieval and recommendation systems, which is used to select the top k items most relevant to the query from a large number of options. According to the calculated similarity results, all non-functional test rules in the pre-built knowledge base are sorted, and the top k rules with the highest similarity are selected, where k≥1. Here, k is a preset positive integer, representing the number of non-functional test rules to be recalled (i.e., selected) from the pre-built knowledge base. The top-k recall algorithm ensures that the selected non-functional test rules are the closest to the current test requirements, which is conducive to generating more relevant and effective non-functional test artifacts and improving the pertinence and efficiency of non-functional testing.
[0061] By vectorizing non-functional test key points and requirement details, using similarity calculation to find the most matching non-functional test rules, and finally using a top-k recall algorithm to select the top k best rules, we ensure that the selected non-functional test rules are both comprehensive and accurate. This approach leverages the advantages of artificial intelligence and machine learning technologies to improve the automation and intelligence level of the non-functional testing process, helping to reduce human error, accelerate the testing process, and ensure test quality and reliability.
[0062] In order to build a rich and structured knowledge base, in some embodiments of the present application, before obtaining the target non-functional test rules that match the above-mentioned non-functional test points and the above-mentioned requirement details from the pre-built knowledge base, and the above-mentioned pre-built knowledge base includes a plurality of non-functional test rules, the above-mentioned method further includes: obtaining historical non-functional test cases and first expert experience data related to the above-mentioned non-functional test, the above-mentioned historical non-functional test cases including historical non-functional test use cases, historical non-functional test plans, historical non-functional test report results and historical non-functional test results; based on the above-mentioned historical non-functional test cases and the above-mentioned first expert experience data, using The above-mentioned pre-trained large language model generates multiple non-functional test rules; a word segmenter is used to split each non-functional test rule into multiple word segmentation units; the multiple word segmentation units are input into a pre-built Transformer model to generate multiple word vectors, and each word vector is updated based on a self-attention mechanism to obtain multiple initial vectors, and the above initial vectors include relevant information between each word segmentation unit of the text information about the non-functional test rules; pooling operation is performed on the multiple initial vectors to generate multiple non-functional test rule vectors, and the above-mentioned pre-built knowledge base composed of multiple non-functional test rule vectors is obtained.
[0063] Specifically, historical non-functional test cases and first-level expert experience data are important components of building a knowledge base. Historical non-functional test cases contain information such as test cases, test plans, test report results, and test results accumulated from previous non-functional tests. First-level expert experience data refers to the expert's professional knowledge and experience in the field of non-functional testing. This data is often contained in documents, reports, training materials written by experts, and records of the experts' actual operations in the project. Historical non-functional test cases and first-level expert experience data are input into a pre-trained large language model, which learns the structure and patterns of this data and generates non-functional test rules. This process requires the design of specific input prompts to guide the model in generating non-functional test rules.
[0064] A tokenizer is a tool used in natural language processing to segment text into smaller words or phrases. The generated non-functional test rules are processed through the tokenizer and split into multiple token units. These token units can be words, phrases, or sentences, depending on the tokenizer configuration. The Transformer model is a deep learning architecture that uses a self-attention mechanism to process sequence data, such as text. The self-attention mechanism allows the model to focus on the associations between various parts of the input sequence, generating a richer vector representation. The tokenized token units are input into the pre-built Transformer model to generate a word vector for each unit. These word vectors are updated through the self-attention mechanism to obtain an initial vector containing the association information between token units.
[0065] Pooling is a common processing method in deep learning models. It is used to extract key information from multiple vectors and generate a comprehensive vector representation. All initial vectors are pooled using average pooling, max pooling, or the more complex attention pooling to generate a non-functional test rule vector. This non-functional test rule vector contains comprehensive information about the entire non-functional test rule and can be used for subsequent similarity calculations and information retrieval.
[0066] A knowledge base is a data structure or database that stores and organizes knowledge data. In this embodiment, the knowledge base stores vector representations of non-functional test rules, facilitating subsequent retrieval and rule selection. All generated non-functional test rule vectors are integrated into a knowledge base, i.e., the pre-built knowledge base described above. This knowledge base can serve as a reference for subsequent non-functional test rule selection and generation.
[0067] The knowledge base in this embodiment is a Retrieval-Augmented Generation (RAG) knowledge base. RAG is a method that combines information retrieval and generative artificial intelligence techniques. It is primarily used to solve the problem of how to integrate external knowledge into the generation process when generating text or content. Specifically, RAG technology allows the model to retrieve and utilize information from a specialized knowledge base when generating content, thereby generating more accurate, detailed, and relevant text.
[0068] Through the above process, a diverse range of non-functional testing rules can be generated from historical cases and expert experience, converted into vector representations, and built into a rich and structured knowledge base. This not only provides a data foundation for subsequent rule selection, but also enhances rule understanding and representation capabilities through pre-trained models and self-attention mechanisms, providing strong support for the automation and intelligentization of software non-functional testing.
[0069] In some embodiments of the present application, according to the above-mentioned target non-functional test rules, non-functional test products are generated using the above-mentioned pre-trained large language model, and the above-mentioned non-functional test products include non-functional test cases, non-functional test plans and non-functional test reports, including: generating prompt words according to the above-mentioned target non-functional test rules and second expert experience data, and the above-mentioned prompt words include: generating test cases from the following five dimensions, the first dimension is to perform a single-interface performance benchmark test on each test interface information, the second dimension is to perform a single-interface load test on each of the above-mentioned test interface information, the third dimension is to perform a single-interface green light test on each of the above-mentioned test interface information, the fourth dimension is to perform a mixed transaction capacity test on all interfaces, and the fifth dimension is to perform a mixed transaction fatigue test on all of the above-mentioned interfaces; the above-mentioned prompt words are input into the above-mentioned pre-trained large language model, and combined with the above-mentioned pre-built knowledge base, the above-mentioned non-functional test products are generated.
[0070] Specifically, in the large language model generation task, prompt words play a role in guiding the model to generate specific types of content, usually containing descriptions, requirements, or contextual information for the generated content. Specific prompt words are created based on the target non-functional test rules and second expert experience data. The second expert experience data focuses more on professional knowledge and experience in the current test project or specific test type. The prompt words require the model to generate test cases along five dimensions. The first dimension is a single-interface performance benchmark test, which means conducting a basic performance evaluation of each test interface, such as response time and throughput. The second dimension is a single-interface load test, which focuses on the performance of the interface under different load conditions, such as the response when the number of concurrent requests is increased. The third dimension is a single-interface green light test, which verifies the functional integrity and stability of the interface under standard conditions. The fourth dimension is a mixed transaction capacity test, which involves the system capacity and performance when multiple interfaces are working simultaneously to ensure that the system can still process requests under high load. The fifth dimension is a mixed transaction fatigue test, which tests the reliability of the system under long-term high-load operation to check for performance degradation or failure issues.
[0071] The prompt words from the five dimensions are fed into a pre-trained large language model. Based on these prompts and its learned knowledge, the model generates non-functional test artifacts, including non-functional test cases, non-functional test plans, and non-functional test reports. These artifacts describe the test procedures, including test steps, parameter settings, and expected results.
[0072] By combining designed prompts with a knowledge base, a pre-trained large language model can generate non-functional test artifacts covering multiple test dimensions. This approach combines the intelligent generation capabilities of AI with the expertise of domain experts, improving the accuracy and professionalism of non-functional test artifacts while also accelerating the preparation and execution of non-functional tests.
[0073] In the complex scenarios of non-functional testing, it is difficult for a single large language model to fully cover all testing requirements and standards. Therefore, this embodiment introduces a multi-model collaborative optimization strategy, which integrates large pre-trained models from multiple different fields to jointly participate in the generation of non-functional test products. For example, for the generation of performance test products, a special performance analysis model and a general large language model can be used in collaboration; for high availability testing, a fault injection model can be introduced to collaborate with the general model. Each model is responsible for generating test content within its professional field, and then integrating and optimizing it to eventually form a comprehensive and sophisticated non-functional test product. The multi-model collaborative optimization strategy can significantly improve the overall quality of non-functional test products, ensuring that the most professional model participates in each test dimension, thereby generating more accurate and comprehensive test plans, test cases and test reports. This method can also improve the resource utilization efficiency of large models, avoid the performance bottleneck of a single model under complex tasks, and make the automation process of non-functional testing more efficient, robust and professional.
[0074] In order to improve the efficiency and accuracy of non-functional test requirement analysis, in some embodiments of the present application, an improvement requirement document of the software to be tested is obtained, and the above-mentioned improvement requirement document is parsed using a pre-trained large language model to extract the non-functional test key points and requirement details, including: obtaining the above-mentioned improvement requirement document of the software to be tested; judging whether the above-mentioned improvement requirement document includes non-functional requirements; if the above-mentioned non-functional requirements are not included in the above-mentioned improvement requirement document, outputting a prompt message and ending the process; if the above-mentioned non-functional requirements are included in the above-mentioned improvement requirement document, inputting the above-mentioned improvement requirement document into the above-mentioned pre-trained large language model, and generating the above-mentioned non-functional test key points and the above-mentioned requirement details according to preset prompt words.
[0075] Specifically, during software development and maintenance, an improvement requirements document is a document that documents improvements or additions to software functionality and performance during maintenance, upgrades, or development. It is crucial for stable software operation and user experience. Improvement requirements documents are key documents used by software development teams and project managers to plan subsequent development work, detailing areas for software improvement. Obtain the improvement requirements document for the software under test from a project management system, document repository, or other source, ensuring it is up-to-date and contains all necessary improvement information.
[0076] Non-functional requirements describe how the software runs, rather than what the software specifically does. They usually involve aspects such as software performance, security, and user experience, and are crucial to ensuring that the software can work properly under various conditions. Through the pre-trained large language model, natural language processing is performed on the improvement requirements document, and the document content is analyzed to determine whether it contains non-functional requirements descriptions. The purpose of this step is to ensure that subsequent non-functional testing covers all relevant areas. If non-functional requirements are not explicitly mentioned in the improvement requirements document, this may mean that the document needs to be further improved, or the scope of non-functional testing needs to be manually defined. If the model determines that there is no non-functional requirement description in the document, a prompt message will be output to inform the user that the document may not contain complete non-functional test requirements, and the user is advised to check or supplement the document, and then end the current non-functional test preparation process.
[0077] If the improvement requirements document contains non-functional requirements, the next step is to use a pre-trained large language model to further parse and extract these requirements in preparation for non-functional testing. Specifically, if the improvement requirements document contains non-functional requirements, it is fed into the pre-trained large language model. Pre-set prompts guide the model in parsing the document, focusing on non-functional descriptions. The model processes the document based on the preset prompts, extracting all relevant non-functional requirements and summarizing them into non-functional test key points and requirement details. These non-functional test key points and requirement details include specific performance testing metrics (such as response time and number of concurrent users), compliance requirements for security testing (such as data encryption and user access control), different environments and configuration requirements for compatibility testing, and duration and load conditions for stability testing. The generated non-functional test key points and requirement details serve as the foundation for subsequent test plans and use case design. A preset prompt might be, "Based on the document content, please extract all requirements related to performance, security, compatibility, and stability, including specific metrics and standards." This prompt guides the model to focus on the non-functional requirements in the document and extract key information.
[0078] By automatically parsing the improved requirements documents for the software under test using a pre-trained large language model, it can intelligently identify information related to non-functional requirements and extract key non-functional test points and requirements details. This not only significantly improves the efficiency and accuracy of requirements analysis, avoiding omissions and misunderstandings that can arise from manual analysis, but also promptly outputs prompt messages and halts unnecessary subsequent testing steps when non-functional requirements are not clearly identified in the requirements document. This saves testing resources, improves the overall intelligence level of test management, and ultimately enhances the efficiency and accuracy of non-functional test requirements analysis.
[0079] In other embodiments of the present application, after generating non-functional test products using the above-mentioned pre-trained large language model according to the above-mentioned target non-functional test rules, the above-mentioned method also includes: if the above-mentioned non-functional test products do not meet the above-mentioned preset standards, adjusting the prompt words input into the above-mentioned pre-trained large language model and the relevant parameters of the above-mentioned pre-trained large language model, and regenerating the above-mentioned non-functional test products, wherein the above-mentioned relevant parameters include learning rate, regularization parameter, context length and generation length.
[0080] Specifically, if the generated non-functional test artifacts do not meet the preset standards, they need to be optimized and regenerated, effectively guiding the non-functional testing process. Prompts are instructions that guide the model to generate specific content types. When non-functional test artifacts do not meet the preset standards, it may be because the prompts fail to accurately express the test requirements or because the model's understanding is biased. By adjusting the prompts, the model can be more accurately guided to generate artifacts that meet the requirements. You can also adjust relevant parameters of the pre-trained large language model, such as the learning rate, regularization parameter, context length, and generation length. The learning rate controls the speed at which the model learns new knowledge. Adjusting the learning rate can help the model learn the intent expressed by the optimized prompts more quickly or more stably. The regularization parameter is used to prevent model overfitting. Adjusting the regularization parameter can prevent the model from being overly dependent on training data and failing to generalize to new generation tasks. The context length is the context within which the model understands the input text. Adjusting the context length can affect the coherence and depth of the generated content. The generation length controls the length of the text generated by the model. Adjusting the generation length ensures that the generated non-functional test artifacts are both detailed and concise, keeping the content concise and focused. The optimized prompt words and adjusted related parameters are input into the pre-trained large language model, and the generation task is re-executed to generate non-functional test artifacts.
[0081] By dynamically adjusting the input prompts and related parameters (such as learning rate, regularization parameter, context length, and generation length) of the pre-trained large language model, the quality and applicability of generated non-functional test artifacts can be significantly improved, ensuring they meet the high standards set by the pre-trained large language model. This mechanism not only enhances the generation flexibility of the pre-trained large language model, but also effectively addresses potential issues that may arise during the initial generation of non-functional test artifacts, such as incomplete content, insufficient accuracy, or non-standard formatting, thereby improving the efficiency of non-functional test preparation and the professionalism of the testing process.
[0082] In some further embodiments of the present application, after generating non-functional test products using the pre-trained large language model according to the target non-functional test rules, the method further includes: obtaining offline evaluation results and online user feedback on the non-functional test products; identifying the non-functional test products with a recall rate lower than a preset threshold based on the offline evaluation results and the online user feedback; converting the positive examples of the non-functional test products with a recall rate lower than the preset threshold into enhanced non-functional test rules to enhance the ability of the pre-trained large language model to generate accurate non-functional test products, and converting the negative examples of the non-functional test products with a recall rate lower than the preset threshold into error-correcting non-functional test rules to guide the pre-trained large language model to identify potential error generation patterns and take measures to prevent the potential errors from occurring again.
[0083] Specifically, offline evaluation results are carefully evaluated by a professional team or an independent third party, typically including a comprehensive review of non-functional test artifacts to assess whether they fully cover non-functional requirements, adhere to correct testing rules, and are easy to understand and execute. Online user feedback comes from test engineers or project managers who actually use these non-functional test artifacts, including feedback on areas where content is unclear, which regulations are not suitable for actual test scenarios, or where test points are missed.
[0084] Recall is a metric that measures whether test artifacts adequately cover all non-functional requirements. A low recall rate indicates that many non-functional requirements are not covered by the tests, or that the generated test artifacts contain omissions and inaccuracies. A standard recall threshold is set; any non-functional test artifact below this threshold is considered a candidate for optimization.
[0085] Non-functional test artifacts that are considered excellent in evaluation and feedback (positive examples) are converted into enhanced non-functional test rules. These rules contain successful elements and patterns and can be used to train the large language model so that it can replicate these successful experiences in future generation. Non-functional test artifacts that reveal problems in evaluation and feedback (negative examples) are converted into error-correcting non-functional test rules. These rules identify common error types and causes, guide the model on how to avoid these errors, and correct potential error generation patterns.
[0086] By converting positive examples into enhanced non-functional test rules, the model can learn and memorize these successful generation patterns, thereby generating more high-quality non-functional test artifacts in the future; by converting negative examples into error-correcting non-functional test rules, the model can identify and self-correct its error tendencies, preventing the same errors from being repeated in future generation processes.
[0087] By collecting offline evaluation results and online user feedback, we identify non-functional test artifacts with low recall rates. These positive and negative examples are then converted into specific training rules, continuously optimizing the generation capabilities of the pre-trained large language model. This mechanism ensures that the generated non-functional test artifacts not only cover all important test points but also significantly improve their accuracy and practicality, reducing the need for manual corrections and enhancing the automation and efficiency of the entire non-functional testing process. Furthermore, it promotes the model's self-learning and evolution, enabling it to adapt to evolving testing requirements and standards.
[0088] In this embodiment, an adaptive model capability enhancement mechanism can also be introduced, which can automatically adjust the training data and training strategy of the large language model according to the specific needs of non-functional testing and the generation effect of non-functional test products. Specifically, by analyzing offline evaluation results and online user feedback, it is possible to identify the weak links of the large language model in specific non-functional testing scenarios, automatically collect high-quality data in related fields to fine-tune the model, and enhance the model's understanding and generation capabilities in these scenarios. In addition, the mechanism can also dynamically adjust training parameters such as learning rate, batch size, etc. according to the generation performance of the large language model to optimize training efficiency and effect. The adaptive model capability enhancement mechanism can continuously improve the intelligent generation capability of the large speech model. Through targeted data enhancement and parameter adjustment, it can significantly improve the generation quality and test efficiency of the test products, reduce manual intervention, and improve the automation and accuracy of the test.
[0089] This embodiment uses the intelligent generation capability based on the large language model, combined with the RAG knowledge base, intelligent agent and other technologies, to input all the required assets for each link of the non-functional test from the beginning. The large language model is responsible for assisting in the generation of test products for all links, and the relevant personnel are responsible for checking and modifying. The output of each link is used as the input of the next link, and finally the entire test task is completed. The large language model intelligent agent refers to an intelligent system that uses a large-scale deep learning model as its core brain and has the ability to independently understand, make decisions and execute complex tasks. Combining the powerful capabilities of the large language model and the autonomy of the intelligent agent, it can provide intelligent services in multiple fields. The embodiment of the present application realizes the intelligence of the non-functional test process by integrating the large language model, RAG retrieval enhancement generation test non-functional test product quality, and intelligent agent technology. From demand analysis to test summary, the large model assists in the generation of each link, manual inspection, adoption or regeneration, which reduces manual operations and improves test efficiency. At the same time, it alleviates the problem of inconsistent test product quality caused by differences in personal experience to a certain extent. By combining prompt template technology based on a large language model with expert experience, we transform the experience of test experts into structured prompt content, extracting test requirements for performance testing, robustness testing, and high-availability testing, and guiding the large language model to generate content that meets testing standards and expert requirements. Integrating the characteristics of non-functional testing within the software system testing process, we combine RAGs with intelligent agents to search knowledge bases and external data, and integrate multiple rounds of intelligent model access to optimize generation quality and mitigate large model hallucinations.
[0090] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the non-functional testing method of the software system of the present application will be described in detail below with reference to specific embodiments.
[0091] This embodiment relates to a specific non-functional testing method for a software system, such as Figure 3As shown, the system automatically parses the modification requirements document for the software under test. The parsing process determines whether the requirements are valid. If they contain no non-functional requirements, a prompt message is output and the task ends. Otherwise, the system searches the knowledge base, accesses the large language model, and applies the agent to generate non-functional test key points and detailed requirements, which are then presented to the user. The user then checks the generated quality. If the quality does not meet the requirements, the large language model prompts and parameters are adjusted, triggering a regeneration. If the quality meets the requirements, the system generates three non-functional test artifacts: non-functional test plans, non-functional test cases, and non-functional test reports. Once these artifacts are manually adopted or regenerated, the test is completed and the process ends. The knowledge base contains the test assets and rules for each stage, which are vectorized and stored in a database. When the user enters a question, the agent's process configuration first searches the knowledge base to obtain relevant information as known information to enhance the quality of model generation. The knowledge base for test plans includes factors to consider for test plan risk analysis and a list of historical test plan risks. The knowledge base for test cases includes tables showing test case types, test case components, design ideas, and example test steps. The test report knowledge base includes factors to consider in test report risk analysis, a list of historical test report risks, and other content. The knowledge base is updated periodically by combining offline evaluation results with online user feedback to analyze scenarios with poor recall rates. Vectorized storage uses a transformer model to generate vectors from the original file, and vectorized retrieval uses a top-k recall method to find the top k most similar content.
[0092] Summarize the non-functional requirements-related parts of the transformation requirements document and utilize the large language model's understanding capabilities and specific prompts to structure and generate several types of test tasks, including performance testing, high availability testing, robustness testing, and interface functional testing. The input required for the non-functional requirements-related parts must include data such as performance indicators, high availability indicators, robustness indicators, and interface functional indicators. The prompts parsed by the large language model describe the thinking behind extracting test tasks, including the following rules for extracting performance test tasks: Based on performance indicators, extract the transaction name, concurrency, response time, etc. required for performance testing; and the following rules for extracting high availability tests: Provide high-availability services as needed, and extract high-availability test tasks based on dimensions such as service switching consistency and service master-slave switching.
[0093] The flow chart of generating non-functional test artifacts based on large language model is as follows Figure 4As shown, the relevant documents of the software to be tested and the historical test assets are stored in the format requirements of the large language model knowledge base. The generation function of each non-functional test product is realized through prompt design and knowledge base retrieval enhancement. The non-functional test products are summarized into a fixed format and output. The non-functional test products are first organized into a knowledge base for intelligent agent application and bare model based on historical assets (historical non-functional test cases) and expert experience data analysis. The non-functional requirements part and expert experience are composed of prompt words according to the preset template to generate test products for the bare model. If the generation quality is unqualified, the following two optimizations can be performed to trigger regeneration: Optimization 1: Optimize the prompt words to trigger the generation of the bare model; Optimization 2: Perform intelligent agent process orchestration to optimize the quality of the generated content by using the positive and negative examples of user feedback as knowledge base content. The default template for the prompt is as follows: You are a performance testing expert, skilled at generating performance test cases. Based on the test interface information provided by the user, you design test cases along the following five dimensions: Dimension 1: Requires a single-interface performance benchmark test for each test interface; Dimension 2: Requires a single-interface load test for each test interface; Dimension 3: Requires a single-interface green light test for each test interface; Dimension 4: Requires a mixed transaction capacity test for all interfaces; Dimension 5: Requires a mixed transaction fatigue test for all interfaces. After the user triggers generation, the generated test artifacts can be displayed on the return page of the user interface. If the generated quality meets the requirements, they will be adopted. If the generation quality does not meet the requirements, the user can click "Record Non-Compliance" to record it in the backend program for subsequent knowledge base additions to improve generation quality. The user can also click the "Regenerate" button to regenerate the artifact.
[0094] The embodiments of the present application also provide a non-functional testing device for a software system. It should be noted that the non-functional testing device for a software system in the embodiments of the present application can be used to execute the non-functional testing method for a software system provided in the embodiments of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and those that have been described will not be repeated here. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0095] The following introduces the non-functional testing device of the software system provided in the embodiment of the present application.
[0096] Figure 5 1 is a structural block diagram of a non-functional testing device for a software system according to an embodiment of the present application. Figure 5As shown, the apparatus includes a parsing unit 10, a first acquisition unit 20, a first generation unit 30, and a non-functional testing unit 40. The parsing unit is configured to acquire an improvement requirement document for the software to be tested and parse the improvement requirement document using a pre-trained large language model to extract non-functional testing key points and requirement details. The improvement requirement document is a document that includes areas where the functions and performance of the software to be tested need to be improved. The first acquisition unit is configured to acquire target non-functional testing rules that match the non-functional testing key points and requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes multiple non-functional testing rules. The first generation unit is configured to generate non-functional test artifacts using the pre-trained large language model based on the target non-functional testing rules. The non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports. The non-functional testing unit is configured to perform non-functional testing on the software to be tested using the non-functional test artifacts if the non-functional test artifacts meet preset standards. The preset standards are based on at least one of the following: integrity, correctness, and executability of the non-functional test artifacts.
[0097] Specifically, after parsing the improvement requirements document, the pre-trained large language model will extract the key points (non-functional test points) and specific requirements (requirements details) related to non-functional features in the improvement requirements document. This process can be seen as converting the improvement requirements of the software to be tested into test specifications, guiding the subsequent non-functional test design and execution. For example, if the improvement requirements document mentions a performance bottleneck, the pre-trained large language model will identify this and parse in detail the various indicators and requirements related to performance testing, ensuring that non-functional testing can comprehensively evaluate the performance improvement effect of the software to be tested.
[0098] Using a pre-trained large language model to parse and improve requirements documents and extract key non-functional test points and requirement details can significantly improve the efficiency, comprehensiveness, and product quality of non-functional testing. This process fully utilizes the natural language understanding and generation capabilities of the large language model, overcoming the limitations of traditional non-functional testing that relies on manual understanding and providing strong data support for automated and intelligent non-functional testing processes.
[0099] A knowledge base is a database system that stores and manages a large amount of structured or semi-structured information. The purpose of constructing a knowledge base is to provide professional knowledge and data support in a field. In an embodiment of the present application, the knowledge base is specifically used in the field of software non-functional testing, including a variety of non-functional testing rules. After the pre-trained large language model extracts the non-functional testing key points and requirement details from the improvement requirement document, the information in the knowledge base is used to match and obtain the most suitable non-functional testing rules, i.e., the target non-functional testing rules. Through the above process, based on the rich information of the improvement requirement document and the knowledge base, non-functional testing rules that match the non-functional testing key points and requirement details can be intelligently and accurately obtained from the pre-built knowledge base, thereby improving the generation quality and pertinence of non-functional testing products, and comprehensively improving the efficiency and quality of software non-functional testing.
[0100] Non-functional test artifacts are generated using a pre-trained large language model. Leveraging its powerful generative capabilities, the model can create text content based on input instructions or rules. Target non-functional test rules are selected from a pre-built knowledge base to best match the non-functional test key points and requirements of the software under test. These rules include specific test types (such as performance testing, security testing, and compatibility testing), test steps, common test methods, and expert experience accumulated from previous testing. The target non-functional test rules are selected based on the parsing results of the improved requirements document, ensuring that the generated non-functional test artifacts are directly targeted at the software's non-functional testing requirements. The pre-trained large language model generates non-functional test cases based on the target non-functional test rules. These non-functional test cases are specific operational instructions or data inputs designed to test the software's performance under specific non-functional metrics. For example, for performance testing, non-functional test cases may include simulating requests under different numbers of concurrent users to evaluate the software's response time and processing capabilities. For security testing, non-functional test cases may involve attempting to inject malicious data or access sensitive resources to test the software's security mechanisms.
[0101] The pre-trained large language model can also generate detailed non-functional test plans. The non-functional test plan is a plan for the entire test process, including test objectives, test environment preparation, test steps, expected results, test resource requirements, etc. The non-functional test plan generated by the pre-trained large language model ensures the systematic and comprehensive nature of the test, and can provide guidance on how to specifically perform the test and how to evaluate the test results. The pre-trained large language model can generate a non-functional test report, which is essentially a predictive report. This non-functional test report will pre-set the expected results, possible risk points, test strategy overview, and test environment configuration during test execution based on the details of the non-functional test cases and non-functional test plans. Such a predictive report can help the test team understand the test plan in advance and be fully prepared for test execution.
[0102] In summary, the pre-trained large language model can generate a series of non-functional test artifacts, including non-functional test cases, non-functional test plans, and non-functional test reports. This process not only significantly improves the efficiency and automation of non-functional testing, but also ensures the quality and professionalism of test artifacts through the large model's generation capabilities. Furthermore, because artifact generation is based on specific target non-functional test rules, the testing process becomes more targeted, better meeting the needs of software non-functional testing and enhancing the intelligence and standardization of the testing process.
[0103] During the software non-functional testing process, the non-functional test artifacts generated by the pre-trained large language model, including non-functional test cases, non-functional test plans, and non-functional test reports, undergo a series of evaluations to ensure they meet pre-set standards. These pre-set standards are based on the completeness, correctness, and enforceability of the non-functional test artifacts to ensure the effectiveness and reliability of the testing process.
[0104] The completeness of non-functional test artifacts refers to whether the non-functional test artifacts cover all necessary test points and requirements, and whether the test plan thoroughly describes each test step and expected results. When evaluating the completeness of non-functional test artifacts, check whether the non-functional test artifacts contain all relevant information such as test types, test cases, and test environment configurations identified based on the improved requirements document and target test rules. For example, performance test artifacts should list in detail all performance indicators that need to be tested, as well as how to design test cases to cover these indicators. The completeness of non-functional test artifacts that meet preset standards can ensure that the test process does not miss important test points, thereby improving the comprehensiveness and effectiveness of the test.
[0105] The accuracy of non-functional test artifacts focuses on whether the information in the artifacts is accurate, whether the test cases truly reflect the key points of non-functional testing, and whether the steps in the test plan are correct. This assessment involves reviewing the artifacts to ensure that the test cases and plans are free of logical errors, incorrect test parameters or test data, and that the test environment is properly configured. Meeting pre-set standards for the accuracy of non-functional test artifacts prevents ineffective or misleading testing due to artifact errors, ensuring the accuracy and reliability of test results.
[0106] The executability of non-functional test artifacts considers their operability and feasibility in the actual test environment. This ensures that the test plan can be executed, that the test cases can be correctly run by the test tool or framework, and that the expected results analysis in the test report can be verified. It is necessary to check whether each step in the test plan is supported by the actual environment, whether the test cases can be successfully executed by the test tool, and whether the analysis in the test report is based on verifiable test data and results. Non-functional test artifact executability that meets preset standards ensures a smooth test process and avoids wasting time and resources due to artifact non-executability.
[0107] Once a non-functional test artifact has passed the pre-defined criteria for completeness, correctness, and executable, it can be used in actual non-functional testing. The generated non-functional test plan is used to prepare the test environment, execute non-functional test cases, and generate a complete non-functional test report based on the actual test results. During this process, the completeness of the non-functional test artifact ensures that all test points are covered, its correctness ensures accurate testing, and its executable ensures the smooth implementation of the test plan, all contributing to the successful execution of the non-functional testing process.
[0108] In the specific implementation process, the first acquisition unit includes a first conversion module, a calculation module, and a determination module. The first conversion module is used to convert the non-functional test key points and the requirement details into vector representations to obtain non-functional test key point vectors and requirement detail vectors; the calculation module is used to calculate the similarity between the non-functional test key point vectors and the requirement detail vectors and the non-functional test rules in the pre-built knowledge base to obtain multiple similarity results; the determination module is used to sort all the non-functional test rules according to the multiple similarity results using the top-k recall algorithm to obtain the top k non-functional test rules with the highest similarity to the non-functional test key point vectors and the requirement detail vectors, and determine the top k non-functional test rules as the target non-functional test rules, where k≥1.
[0109] Specifically, a pre-trained language model such as Transformer can be used to convert the non-functional test key points and requirement details into vector representations to obtain non-functional test key point vectors and requirement detail vectors. Through vectorization, non-functional test key points and requirement details can be understood and processed by the computer system, providing a basis for subsequent similarity calculations. Similarity calculations can measure the degree of similarity between two vectors and can be implemented through a variety of distance or similarity measurement methods, such as cosine similarity and Euclidean distance. Similarity calculations are performed on the non-functional test key point vectors and requirement detail vectors respectively with each non-functional test rule vector in the pre-built knowledge base. For example, using cosine similarity, the cosine value of the angle between the two vectors is calculated. The closer the value is to 1, the more similar the two are. The results of the similarity calculation can help identify which non-functional test rules are most compatible with the current non-functional test key point vectors and requirement details, thereby optimizing the selection of non-functional test rules.
[0110] After obtaining the similarity results, the top-k recall algorithm is used to sort all non-functional test rules. The top-k recall algorithm is an algorithm used in the field of information retrieval and recommendation systems, which is used to select the top k items most relevant to the query from a large number of options. According to the calculated similarity results, all non-functional test rules in the pre-built knowledge base are sorted, and the top k rules with the highest similarity are selected, where k≥1. Here, k is a preset positive integer, representing the number of non-functional test rules to be recalled (i.e., selected) from the pre-built knowledge base. The top-k recall algorithm ensures that the selected non-functional test rules are the closest to the current test requirements, which is conducive to generating more relevant and effective non-functional test artifacts and improving the pertinence and efficiency of non-functional testing.
[0111] By vectorizing non-functional test key points and requirement details, using similarity calculation to find the most matching non-functional test rules, and finally using a top-k recall algorithm to select the top k best rules, we ensure that the selected non-functional test rules are both comprehensive and accurate. This approach leverages the advantages of artificial intelligence and machine learning technologies to improve the automation and intelligence level of the non-functional testing process, helping to reduce human error, accelerate the testing process, and ensure test quality and reliability.
[0112] In order to build a rich and structured knowledge base, in some embodiments of the present application, the above-mentioned device also includes a second acquisition unit, a second generation unit, a splitting unit, an input unit and a pooling operation unit. The second acquisition unit is used to obtain the target non-functional test rules that match the above-mentioned non-functional test points and the above-mentioned requirement details from the pre-built knowledge base, and the above-mentioned pre-built knowledge base includes a variety of non-functional test rules. The historical non-functional test cases include historical non-functional test cases, historical non-functional test plans, historical non-functional test report results and historical non-functional test results; the second generation unit is used to generate multiple non-functional test cases based on the above-mentioned historical non-functional test cases and the above-mentioned first expert experience data using the above-mentioned pre-trained large language model. Functional testing rules; the splitting unit is used to split each of the above non-functional testing rules into multiple word segmentation units using a word segmenter; the input unit is used to input the multiple word segmentation units into a pre-built Transformer model to generate multiple word vectors, and update the multiple word vectors based on the self-attention mechanism to obtain multiple initial vectors, and the above initial vectors include relevant information between the multiple word segmentation units of the text information about the above non-functional testing rules; the pooling operation unit is used to perform pooling operations on the multiple initial vectors to generate multiple non-functional testing rule vectors, and obtain the above pre-built knowledge base composed of the multiple non-functional testing rule vectors.
[0113] Specifically, historical non-functional test cases and first-level expert experience data are important components of building a knowledge base. Historical non-functional test cases contain information such as test cases, test plans, test report results, and test results accumulated from previous non-functional tests. First-level expert experience data refers to the expert's professional knowledge and experience in the field of non-functional testing. This data is often contained in documents, reports, training materials written by experts, and records of the experts' actual operations in the project. Historical non-functional test cases and first-level expert experience data are input into a pre-trained large language model, which learns the structure and patterns of this data and generates non-functional test rules. This process requires the design of specific input prompts to guide the model in generating non-functional test rules.
[0114] A tokenizer is a tool used in natural language processing to segment text into smaller words or phrases. The generated non-functional test rules are processed through the tokenizer and split into multiple token units. These token units can be words, phrases, or sentences, depending on the tokenizer configuration. The Transformer model is a deep learning architecture that uses a self-attention mechanism to process sequence data, such as text. The self-attention mechanism allows the model to focus on the associations between various parts of the input sequence, generating a richer vector representation. The tokenized token units are input into the pre-built Transformer model to generate a word vector for each unit. These word vectors are updated through the self-attention mechanism to obtain an initial vector containing the association information between token units.
[0115] Pooling is a common processing method in deep learning models. It is used to extract key information from multiple vectors and generate a comprehensive vector representation. All initial vectors are pooled using average pooling, max pooling, or the more complex attention pooling to generate a non-functional test rule vector. This non-functional test rule vector contains comprehensive information about the entire non-functional test rule and can be used for subsequent similarity calculations and information retrieval.
[0116] A knowledge base is a data structure or database that stores and organizes knowledge data. In this embodiment, the knowledge base stores vector representations of non-functional test rules, facilitating subsequent retrieval and rule selection. All generated non-functional test rule vectors are integrated into a knowledge base, i.e., the pre-built knowledge base described above. This knowledge base can serve as a reference for subsequent non-functional test rule selection and generation.
[0117] Through the above process, a diverse range of non-functional testing rules can be generated from historical cases and expert experience, converted into vector representations, and built into a rich and structured knowledge base. This not only provides a data foundation for subsequent rule selection, but also enhances rule understanding and representation capabilities through pre-trained models and self-attention mechanisms, providing strong support for the automation and intelligentization of software non-functional testing.
[0118] In some embodiments of the present application, the first generation unit includes a first generation module and a second generation module. The first generation module is used to generate prompt words based on the target non-functional test rules and the second expert experience data, and the prompt words include: generating test cases from the following five dimensions: the first dimension is to perform a single-interface performance benchmark test on each test interface information, the second dimension is to perform a single-interface load test on each of the above test interface information, the third dimension is to perform a single-interface green light test on each of the above test interface information, the fourth dimension is to perform a mixed transaction capacity test on all interfaces, and the fifth dimension is to perform a mixed transaction fatigue test on all of the above interfaces; the second generation module is used to input the prompt words into the pre-trained large language model and generate the non-functional test artifacts in combination with the pre-built knowledge base.
[0119] Specifically, in the large language model generation task, prompt words play a role in guiding the model to generate specific types of content, usually containing descriptions, requirements, or contextual information for the generated content. Specific prompt words are created based on the target non-functional test rules and second expert experience data. The second expert experience data focuses more on professional knowledge and experience in the current test project or specific test type. The prompt words require the model to generate test cases along five dimensions. The first dimension is a single-interface performance benchmark test, which means conducting a basic performance evaluation of each test interface, such as response time and throughput. The second dimension is a single-interface load test, which focuses on the performance of the interface under different load conditions, such as the response when the number of concurrent requests is increased. The third dimension is a single-interface green light test, which verifies the functional integrity and stability of the interface under standard conditions. The fourth dimension is a mixed transaction capacity test, which involves the system capacity and performance when multiple interfaces are working simultaneously to ensure that the system can still process requests under high load. The fifth dimension is a mixed transaction fatigue test, which tests the reliability of the system under long-term high-load operation to check for performance degradation or failure issues.
[0120] The prompt words from the five dimensions are fed into a pre-trained large language model. Based on these prompts and its learned knowledge, the model generates non-functional test artifacts, including non-functional test cases, non-functional test plans, and non-functional test reports. These artifacts describe the test procedures, including test steps, parameter settings, and expected results.
[0121] By combining designed prompts with a knowledge base, a pre-trained large language model can generate non-functional test artifacts covering multiple test dimensions. This approach combines the intelligent generation capabilities of AI with the expertise of domain experts, improving the accuracy and professionalism of non-functional test artifacts while also accelerating the preparation and execution of non-functional tests.
[0122] To improve the efficiency and accuracy of non-functional test requirement analysis, in some embodiments of the present application, the parsing unit includes: an acquisition module, an output module, and a third generation module. The acquisition module is used to acquire the improvement requirement document of the software to be tested; determine whether the improvement requirement document includes non-functional requirements; the output module is used to output a prompt message and terminate the process if the improvement requirement document does not include the non-functional requirements; and the third generation module is used to input the improvement requirement document into the pre-trained large language model if the improvement requirement document includes the non-functional requirements, and generate the non-functional test key points and the requirement details based on preset prompt words.
[0123] Specifically, during software development and maintenance, an improvement requirements document is a document that documents improvements or additions to software functionality and performance during maintenance, upgrades, or development. It is crucial for stable software operation and user experience. Improvement requirements documents are key documents used by software development teams and project managers to plan subsequent development work, detailing areas for software improvement. Obtain the improvement requirements document for the software under test from a project management system, document repository, or other source, ensuring it is up-to-date and contains all necessary improvement information.
[0124] Non-functional requirements describe how the software runs, rather than what the software specifically does. They usually involve aspects such as software performance, security, and user experience, and are crucial to ensuring that the software can work properly under various conditions. Through the pre-trained large language model, natural language processing is performed on the improvement requirements document, and the document content is analyzed to determine whether it contains non-functional requirements descriptions. The purpose of this step is to ensure that subsequent non-functional testing covers all relevant areas. If non-functional requirements are not explicitly mentioned in the improvement requirements document, this may mean that the document needs to be further improved, or the scope of non-functional testing needs to be manually defined. If the model determines that there is no non-functional requirement description in the document, a prompt message will be output to inform the user that the document may not contain complete non-functional test requirements, and the user is advised to check or supplement the document, and then end the current non-functional test preparation process.
[0125] If the improvement requirements document contains non-functional requirements, the next step is to use a pre-trained large language model to further parse and extract these requirements in preparation for non-functional testing. Specifically, if the improvement requirements document contains non-functional requirements, it is fed into the pre-trained large language model. Pre-set prompts guide the model in parsing the document, focusing on non-functional descriptions. The model processes the document based on the preset prompts, extracting all relevant non-functional requirements and summarizing them into non-functional test key points and requirement details. These non-functional test key points and requirement details include specific performance testing metrics (such as response time and number of concurrent users), compliance requirements for security testing (such as data encryption and user access control), different environments and configuration requirements for compatibility testing, and duration and load conditions for stability testing. The generated non-functional test key points and requirement details serve as the foundation for subsequent test plans and use case design. A preset prompt might be, "Based on the document content, please extract all requirements related to performance, security, compatibility, and stability, including specific metrics and standards." This prompt guides the model to focus on the non-functional requirements in the document and extract key information.
[0126] By automatically parsing the improved requirements documents for the software under test using a pre-trained large language model, it can intelligently identify information related to non-functional requirements and extract key non-functional test points and requirements details. This not only significantly improves the efficiency and accuracy of requirements analysis, avoiding omissions and misunderstandings that can arise from manual analysis, but also promptly outputs prompt messages and halts unnecessary subsequent testing steps when non-functional requirements are not clearly identified in the requirements document. This saves testing resources, improves the overall intelligence level of test management, and ultimately enhances the efficiency and accuracy of non-functional test requirements analysis.
[0127] In some other embodiments of the present application, the above-mentioned device also includes an adjustment unit for adjusting the prompt words input into the above-mentioned pre-trained large language model and relevant parameters of the above-mentioned pre-trained large language model to regenerate the above-mentioned non-functional test product after generating the non-functional test product according to the above-mentioned target non-functional test rules using the above-mentioned pre-trained large language model if the above-mentioned non-functional test product does not meet the above-mentioned preset standards, wherein the above-mentioned relevant parameters include learning rate, regularization parameter, context length and generation length.
[0128] Specifically, if the generated non-functional test artifacts do not meet the preset standards, they need to be optimized and regenerated, effectively guiding the non-functional testing process. Prompts are instructions that guide the model to generate specific content types. When non-functional test artifacts do not meet the preset standards, it may be because the prompts fail to accurately express the test requirements or because the model's understanding is biased. By adjusting the prompts, the model can be more accurately guided to generate artifacts that meet the requirements. You can also adjust relevant parameters of the pre-trained large language model, such as the learning rate, regularization parameter, context length, and generation length. The learning rate controls the speed at which the model learns new knowledge. Adjusting the learning rate can help the model learn the intent expressed by the optimized prompts more quickly or more stably. The regularization parameter is used to prevent model overfitting. Adjusting the regularization parameter can prevent the model from being overly dependent on training data and failing to generalize to new generation tasks. The context length is the context within which the model understands the input text. Adjusting the context length can affect the coherence and depth of the generated content. The generation length controls the length of the text generated by the model. Adjusting the generation length ensures that the generated non-functional test artifacts are both detailed and concise, keeping the content concise and focused. The optimized prompt words and adjusted related parameters are input into the pre-trained large language model, and the generation task is re-executed to generate non-functional test artifacts.
[0129] By dynamically adjusting the input prompts and related parameters (such as learning rate, regularization parameter, context length, and generation length) of the pre-trained large language model, the quality and applicability of generated non-functional test artifacts can be significantly improved, ensuring they meet the high standards set by the pre-trained large language model. This mechanism not only enhances the generation flexibility of the pre-trained large language model, but also effectively addresses potential issues that may arise during the initial generation of non-functional test artifacts, such as incomplete content, insufficient accuracy, or non-standard formatting, thereby improving the efficiency of non-functional test preparation and the professionalism of the testing process.
[0130] In some further embodiments of the present application, the apparatus further includes a third acquisition unit, an identification unit, and a conversion unit. The third acquisition unit is configured to obtain offline evaluation results and online user feedback on the non-functional test product after the pre-trained large language model is used to generate the non-functional test product according to the target non-functional test rule; the identification unit is configured to identify the non-functional test product with a recall rate lower than a preset threshold based on the offline evaluation results and the online user feedback; the conversion unit is configured to convert the positive examples of the non-functional test product with a recall rate lower than the preset threshold into enhanced non-functional test rules to enhance the ability of the pre-trained large language model to generate accurate non-functional test products, and to convert the negative examples of the non-functional test product with a recall rate lower than the preset threshold into error-correcting non-functional test rules to guide the pre-trained large language model to identify potential error generation patterns and take measures to prevent the potential errors from recurring.
[0131] Specifically, offline evaluation results are carefully evaluated by a professional team or an independent third party, typically including a comprehensive review of non-functional test artifacts to assess whether they fully cover non-functional requirements, adhere to correct testing rules, and are easy to understand and execute. Online user feedback comes from test engineers or project managers who actually use these non-functional test artifacts, including feedback on areas where content is unclear, which regulations are not suitable for actual test scenarios, or where test points are missed.
[0132] Recall is a metric that measures whether test artifacts adequately cover all non-functional requirements. A low recall rate indicates that many non-functional requirements are not covered by the tests, or that the generated test artifacts contain omissions and inaccuracies. A standard recall threshold is set; any non-functional test artifact below this threshold is considered a candidate for optimization.
[0133] Non-functional test artifacts that are considered excellent in evaluation and feedback (positive examples) are converted into enhanced non-functional test rules. These rules contain successful elements and patterns and can be used to train the large language model so that it can replicate these successful experiences in future generation. Non-functional test artifacts that reveal problems in evaluation and feedback (negative examples) are converted into error-correcting non-functional test rules. These rules identify common error types and causes, guide the model on how to avoid these errors, and correct potential error generation patterns.
[0134] By converting positive examples into enhanced non-functional test rules, the model can learn and memorize these successful generation patterns, thereby generating more high-quality non-functional test artifacts in the future; by converting negative examples into error-correcting non-functional test rules, the model can identify and self-correct its error tendencies, preventing the same errors from being repeated in future generation processes.
[0135] By collecting offline evaluation results and online user feedback, we identify non-functional test artifacts with low recall rates. These positive and negative examples are then converted into specific training rules, continuously optimizing the generation capabilities of the pre-trained large language model. This mechanism ensures that the generated non-functional test artifacts not only cover all important test points but also significantly improve their accuracy and practicality, reducing the need for manual corrections and enhancing the automation and efficiency of the entire non-functional testing process. Furthermore, it promotes the model's self-learning and evolution, enabling it to adapt to evolving testing requirements and standards.
[0136] The non-functional testing device for the software system includes a processor and a memory. The parsing unit, first acquisition unit, first generation unit, and non-functional testing unit are all stored as program units in the memory. The processor executes the program units stored in the memory to implement the corresponding functions. The modules are all located in the same processor; alternatively, the modules can be located in different processors in any combination.
[0137] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0138] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute the non-functional testing method of the software system.
[0139] An embodiment of the present invention provides an electronic device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps of the non-functional testing method of the software system are implemented.
[0140] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of initializing the non-functional testing method of the above-mentioned software system.
[0141] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0142] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0143] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0144] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0146] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0147] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0148] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0149] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0150] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A non-functional testing method for a software system, characterized in that: The non-functional testing method includes multiple types of non-functional testing, each type of non-functional testing includes corresponding non-functional testing points and requirement details, including: Obtain an improvement requirement document for the software to be tested, and parse the improvement requirement document using a pre-trained large language model to extract non-functional test points and requirement details. The improvement requirement document refers to a document that includes areas where the functions and performance of the software to be tested need to be improved. Acquire target non-functional test rules that match the non-functional test key points and the requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes a plurality of non-functional test rules; Generate non-functional test artifacts using the pre-trained large language model according to the target non-functional test rules, wherein the non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports; If the non-functional test artifact meets a preset standard, the non-functional test artifact is used to perform non-functional testing on the software to be tested, wherein the preset standard is formulated based on at least one of the following: integrity, correctness and executableness of the non-functional test artifact.
2. The method according to claim 1, characterized in that The non-functional test rules are in vector form, and target non-functional test rules that match the non-functional test points and the requirement details are obtained from a pre-built knowledge base. The pre-built knowledge base includes multiple non-functional test rules, including: Converting the non-functional test key points and the requirement details into vector representations to obtain a non-functional test key point vector and a requirement detail vector; Calculating similarities between the non-functional test key point vector and the requirement detail vector and each of the non-functional test rules in the pre-built knowledge base, to obtain a plurality of similarity results; Based on multiple similarity results, a top-k recall algorithm is used to sort all the non-functional test rules to obtain the top k non-functional test rules with the highest similarity to the non-functional test key point vector and the requirement detail vector, and the top k non-functional test rules are determined as the target non-functional test rules, where k≥1.
3. The method according to claim 1, characterized in that Before obtaining a target non-functional test rule that matches the non-functional test key points and the requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes a plurality of non-functional test rules, the method further includes: Acquire historical non-functional test cases and first expert experience data related to the non-functional test, wherein the historical non-functional test cases include historical non-functional test use cases, historical non-functional test plans, historical non-functional test report results, and historical non-functional test results; Generating a plurality of non-functional test rules using the pre-trained large language model according to the historical non-functional test cases and the first expert experience data; Using a word segmenter to split each of the non-functional test rules into multiple word segmentation units; Inputting the plurality of word segmentation units into a pre-built Transformer model to generate a plurality of word vectors, and updating each of the word vectors based on a self-attention mechanism to obtain a plurality of initial vectors, wherein the initial vectors include relevant information between each of the word segmentation units regarding the text information of the non-functional test rule; A pooling operation is performed on the multiple initial vectors to generate multiple non-functional test rule vectors, thereby obtaining the pre-built knowledge base composed of the multiple non-functional test rule vectors.
4. The method according to claim 1, wherein According to the target non-functional test rules, non-functional test artifacts are generated using the pre-trained large language model. The non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports, including: Generate prompt words based on the target non-functional test rules and the second expert experience data, the prompt words including: generating test cases from the following five dimensions: a first dimension is performing a single-interface performance benchmark test on each test interface information; a second dimension is performing a single-interface load test on each test interface information; a third dimension is performing a single-interface green light test on each test interface information; a fourth dimension is performing a mixed transaction capacity test on all interfaces; and a fifth dimension is performing a mixed transaction fatigue test on all interfaces; The prompt words are input into the pre-trained large language model and combined with the pre-built knowledge base to generate the non-functional test artifact.
5. The method according to claim 1, wherein Obtain the improvement requirement document of the software to be tested, and use the pre-trained large language model to parse the improvement requirement document to extract non-functional test points and requirement details, including: Obtaining the improvement requirement document of the software to be tested; Determining whether the improvement requirement document includes non-functional requirements; If the non-functional requirement is not included in the improvement requirement document, output a prompt message and end the process; In the case where the improvement requirement document includes the non-functional requirement, the improvement requirement document is input into the pre-trained large language model, and the non-functional test points and the requirement details are generated according to preset prompt words.
6. The method according to claim 1, characterized in that After generating non-functional test artifacts using the pre-trained large language model according to the target non-functional test rules, the method further includes: If the non-functional test product does not meet the preset standard, the prompt words input into the pre-trained large language model and the relevant parameters of the pre-trained large language model are adjusted to regenerate the non-functional test product. The relevant parameters include learning rate, regularization parameter, context length and generation length.
7. The method according to claim 1, characterized in that After generating non-functional test artifacts using the pre-trained large language model according to the target non-functional test rules, the method further includes: Obtaining offline evaluation results and online user feedback on the non-functional test product; Identifying the non-functional test products having a recall rate lower than a preset threshold according to the offline evaluation results and the online user feedback; The positive examples of the non-functional test products with a recall rate lower than the preset threshold are converted into enhanced non-functional test rules to enhance the ability of the pre-trained large language model to generate accurate non-functional test products, and the negative examples of the non-functional test products with a recall rate lower than the preset threshold are converted into error-correcting non-functional test rules to guide the pre-trained large language model to identify potential error generation patterns and take measures to prevent the potential errors from occurring again.
8. A non-functional testing device for a software system, characterized in that: include: A parsing unit is configured to obtain an improvement requirement document of the software to be tested, and parse the improvement requirement document using a pre-trained large language model to extract non-functional test points and requirement details. The improvement requirement document refers to a document that includes areas where the functions and performance of the software to be tested need to be improved. A first acquisition unit is configured to acquire a target non-functional test rule that matches the non-functional test key points and the requirement details from a pre-built knowledge base, wherein the pre-built knowledge base includes a plurality of non-functional test rules; A first generating unit is configured to generate non-functional test artifacts using the pre-trained large language model according to the target non-functional test rules, wherein the non-functional test artifacts include non-functional test cases, non-functional test plans, and non-functional test reports; A non-functional testing unit is configured to perform non-functional testing on the software to be tested using the non-functional testing artifact if the non-functional testing artifact meets a preset standard, wherein the preset standard is formulated based on at least one of the following: integrity, correctness, and executability of the non-functional testing artifact.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the non-functional testing method for the software system according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a non-functional testing method for executing the software system according to any one of claims 1 to 7.
Citation Information
Cited By
Test case generation method and device
CN120973690A