An artificial intelligence-based method for evaluating the compliance of software product requirements implementation

Through artificial intelligence-based methods, offline large models are used to generate test cases and operating instructions, and the compliance of IPTV software products with requirements is automatically evaluated. This solves the problems of low efficiency and high cost of traditional acceptance methods and realizes efficient and reliable acceptance assistance.

CN118779238BActive Publication Date: 2025-10-03PACO VIDEO TECH (HANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410937226.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-10-03
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

In the IPTV field, traditional software product acceptance methods are inefficient and costly. They cannot effectively verify whether the product has achieved the design requirements and there is a risk of vulnerabilities. Existing automatic dialing technology and manual acceptance solutions are both insufficient.

Method used

An AI-based approach is used to train the offline large-scale ChatGLM-6B model to generate test cases, operation instructions, and expected result diagrams, automatically evaluating the compliance of software product requirements. The first model is used to generate a complete set of test cases and operation instruction sets, and the second model automatically calls the software product for screenshot evaluation.

Benefits of technology

It improves the coverage and efficiency of software product acceptance, reduces manpower input costs, ensures acceptance quality, reduces loopholes and omissions, and provides reliable acceptance assistance tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118779238B_ABST
    Figure CN118779238B_ABST
Patent Text Reader

Abstract

The present application provides a method for evaluating the compliance of software product requirements based on artificial intelligence, including: obtaining the product requirement specification, product page UI design, product page interaction design, prior test case set and software product of the project to be evaluated, and inputting them into the trained first model, outputting the completed test case set, operation instruction set and expected result atlas through the first model; inputting the completed test case set, operation instruction set and expected result atlas into the second model and calling the software product, calling the corresponding operation instruction in the operation instruction set based on the completed test case set through the second model to run the software product and automatically take screenshots to form an operation result atlas, and generating the compliance of the software product of the project to be evaluated based on the expected result atlas and the operation result atlas. By means of artificial intelligence, the compliance of the requirements is evaluated, the acceptance coverage and acceptance efficiency are effectively improved, and reliable acceptance assistance is provided for acceptance personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to an artificial intelligence-based method for evaluating the compliance of software product requirements implementation. Background Art

[0002] In the IPTV industry, every EPG (Electronic Program Guide) product development process requires first designing a UI prototype, then designing the interaction design based on the prototype, and finally writing a product requirements document. Testers then write test cases based on these three documents for testing. After R&D and testing, the released product undergoes product acceptance by both internal and client product managers. This entire acceptance cycle is extremely time-consuming, and even after issues are fixed, the entire process must be repeated (because it's uncertain whether the fix will affect other features that have already been accepted). Furthermore, due to the limited energy and investment of product managers, some features and product vulnerabilities may be missed. IPTV is played on set-top boxes, which imposes stringent requirements for vulnerability risk management.

[0003] There are two traditional solutions:

[0004] 1. Using automatic dialing technology requires purchasing specialized hardware and software services, which is costly. Furthermore, automatic dialing technology can only automatically detect whether the EPG is missing content or if any content cannot be played, but it cannot verify whether the product meets the required design requirements or achieves the expected results.

[0005] 2. Build an acceptance team and increase manpower to conduct comprehensive acceptance testing. Multiple testers and product managers are assigned to conduct multiple rounds of acceptance verification, and even multiple customer service representatives may be assigned to assist. This approach is extremely inefficient, with limited acceptance coverage and unreliable acceptance quality. It also carries high costs and a long lead time.

[0006] Therefore, how to provide an intelligent acceptance assistance solution for software products and effectively improve the acceptance efficiency of software products with a lower cost investment is a problem that needs to be solved at present. Summary of the Invention

[0007] The purpose of the embodiments of the present application is to provide a software product requirement implementation compliance assessment method based on artificial intelligence, so as to intelligently generate test cases, operation instructions and expected result diagrams based on product requirement specifications, product page UI design, and product page interaction design through artificial intelligence, and use the operation instructions to run the software to automatically take screenshots. After obtaining the operation result diagram, it is combined with the expected result diagram to perform an intelligent assessment of the compliance with the requirements, thereby effectively improving the acceptance coverage and acceptance efficiency, and providing reliable acceptance assistance for acceptance personnel.

[0008] In order to achieve the above objectives, the embodiments of the present application are implemented in the following manner:

[0009] In a first aspect, an embodiment of the present application provides a method for evaluating the compliance of software product requirements based on artificial intelligence, comprising: obtaining a product requirement specification, a product page UI design, a product page interaction design, a priori test case set, and a software product of a project to be evaluated, wherein the priori test case set is an empty set or a non-empty set, and the test cases in the priori test case set are generated by product managers and / or testers for testing some functions of the software product; inputting the product requirement specification, the product page UI design, the product page interaction design, and the priori test case set into a trained first model, and outputting a completed test case set, an operation instruction set, and An expected result atlas, wherein the completion test case set includes a priori test case set, and each test case in the completion test case set corresponds to an operation instruction chain in the operation instruction set and an expected result graph in the expected result atlas, and each operation instruction chain includes at least one operation instruction; the completion test case set, the operation instruction set and the expected result atlas are input into the second model and the software product is called, and the second model calls the corresponding operation instruction in the operation instruction set based on the completion test case set to run the software product and automatically take a screenshot to form an operation result atlas, and based on the expected result atlas and the operation result atlas, the requirement implementation compliance of the software product of the project to be evaluated is generated.

[0010] In combination with the first aspect, in a first possible implementation of the first aspect, the first model is obtained by private training using the offline large model ChatGLM-6B. The training process of the first model is as follows: a historical project data set is obtained, the historical project data set contains multiple historical project data, each historical project data contains a software product, a product requirement specification, a product page UI design, a product page interaction design, a test case subset, an operation instruction subset, and an expected result graph subset. The test case subset contains N test cases, the operation instruction subset contains N operation instruction chains, each operation instruction chain contains at least one operation instruction, and the expected result graph subset contains N expected results. The historical project data set is divided into a training set and a validation set; the product requirement specifications, product page UI design, product page interaction design, a subset of test cases, a subset of operation instructions, and a subset of expected result diagrams of the historical project data in the training set are imported into the offline large model ChatGLM-6B for training; the product requirement specifications, product page UI design, product page interaction design, a subset of test cases, a subset of operation instructions, and a subset of expected result diagrams of the historical project data in the validation set are imported into the trained offline large model ChatGLM-6B for verification and optimization, and the optimized offline large model ChatGLM-6B is obtained as the first model.

[0011] In combination with the first possible implementation method of the first aspect, in the second possible implementation method of the first aspect, the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result diagram subset of the historical project data in the training set are imported into the offline large model ChatGLM-6B for training, including: importing the product page UI design of the historical project data in the training set into the offline large model ChatGLM-6B for private training, which is used to analyze the various elements in the product page UI design and obtain sub-model modelA; importing the product page interaction design and operation instruction subset of the historical project data in the training set into the offline large model ChatGLM-6B for private training, which is used to analyze the correspondence between the operation instructions and the interactions in the product page interaction design and obtain sub-model modelB; importing the product requirement specifications of the historical project data in the training set into the offline large model ChatGLM-6B for private training, which is used to analyze product requirements and obtain sub-model modelC; importing the historical project data in the training set into the offline large model ChatGLM-6B for private training The test case subset is imported into the offline large model ChatGLM-6B for private training, which is used to analyze the test cases and obtain the sub-model modelD; the association between the offline large model ChatGLM-6B and the sub-model modelA, sub-model modelB, sub-model modelC and sub-model modelD is established, and the product requirement specification, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the training set are imported for association training to obtain the integrated model modelbase, and the generative test case corresponding to each requirement in the product requirement specification of the historical project data, the generative operation instruction chain corresponding to each generative test case and the generative expected result graph are generated through the integrated model modelbase, and based on the test case subset, operation instruction subset and expected result graph subset of the historical project data, each generative test case of the historical project data, the generative operation instruction chain corresponding to each generative test case and the generative expected result graph are fed back and updated.

[0012] In combination with the second possible implementation method of the first aspect, in the third possible implementation method of the first aspect, the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the verification set are imported into the trained offline large model ChatGLM-6B for verification and optimization, including: importing the product requirement specifications, product page UI design, product page interaction design of the historical project data in the verification set into the integrated model modelbase; for each historical project data in the verification set, the integrated model modelbase generates a generative test case corresponding to each requirement in the product requirement specification, a generative operation instruction chain corresponding to each generative test case and a generative expected result graph based on the product requirement specifications, product page UI design and product page interaction design of the historical project data; for the generative test case corresponding to each requirement: obtaining an accuracy score group of the generative test case, and determining based on the accuracy score group Obtain the accuracy score of this generative test case, and then verify and optimize the integrated model modelbase based on the accuracy score of this generative test case and the test case corresponding to this requirement in the test case set, wherein the accuracy score group includes the accuracy scores of this generative test case by multiple testers and / or product managers; for the generative operation instruction chain corresponding to each generative test case: obtain the accuracy score group of the generative operation instruction chain, and determine the accuracy score of this generative operation instruction chain based on the accuracy score group, and then verify and optimize the integrated model modelbase based on the accuracy score of this generative operation instruction chain and the operation instruction chain of the corresponding test case, wherein the accuracy score group includes the accuracy scores of this generative operation instruction chain by multiple testers and / or product managers; for the generative expected result graph corresponding to each generative test case: verify and optimize the integrated model modelbase based on the generative expected result graph and the corresponding expected result graph in the expected result graph subset.

[0013] In combination with the third possible implementation manner of the first aspect, in a fourth possible implementation manner of the first aspect, determining the accuracy score of the generative test case based on the accuracy score group includes: determining the source role and score of each accuracy score in the accuracy score group of the generative test case, wherein the source role is a tester or a product manager; and calculating the accuracy score of the generative test case based on the source role and score of each accuracy score in the following manner:

[0014]

[0015] Among them, R is the accuracy score of the generated test case, is the mean accuracy score of the testers for this generated test case, is the mean accuracy score of the product manager for this generative test case, n T The number of testers who rated the accuracy of this generated test case, n P The number of product managers who rated the accuracy of this generative test case, is the accuracy score of the generative test case given by the i-th tester, The accuracy score of the generative test case given by the j-th product manager.

[0016] In conjunction with the third possible implementation manner of the first aspect, in a fifth possible implementation manner of the first aspect, determining the accuracy score of the generative operation instruction chain based on the accuracy score group includes: calculating the accuracy score of the generative operation instruction chain based on the score of each accuracy score in the accuracy score group of the generative operation instruction chain using the following method:

[0017]

[0018] Among them, S is the accuracy score of the generated operation instruction chain, n is the total number of people who score the accuracy of this generated operation instruction chain, and the people are testers or product managers. i The accuracy score of the generated operation instruction chain for the i-th person.

[0019] In combination with the first possible implementation method of the first aspect, in the sixth possible implementation method of the first aspect, the second model is obtained by privatization training using the offline large model ChatGLM-6B, and the training process of the second model is as follows: based on the software product of each historical project data in the historical project data set, a corresponding software call path is generated, and the software call path is imported into the offline large model ChatGLM-6B; for each historical project data, a subset of test cases, a subset of operation instructions, and a subset of expected result diagrams of the historical project data are imported into the offline large model ChatGLM-6B, and the software call path is called by the offline large model ChatGLM-6B to run the software product, and the test case subset, the operation instruction subset, and the expected result diagram subset are imported into the offline large model ChatGLM-6B. For each test case in the trial example set, control the software product to execute the operation instruction chain of the test case, take a screenshot of the software product interface after executing the operation instruction chain, and obtain the operation result graph corresponding to the test case. Based on the operation result graph and the corresponding expected result graph, determine the test requirement compliance score of this operation result graph, and obtain the calibration requirement compliance score group of this operation result graph, and determine the calibration requirement compliance score of this operation result graph based on the calibration requirement compliance score group. Based on the calibration requirement compliance score and the test requirement compliance score of this operation result graph, feedback and update the offline large model ChatGLM-6B, and finally obtain the trained second model.

[0020] In combination with the sixth possible implementation manner of the first aspect, in a seventh possible implementation manner of the first aspect, determining the calibration requirement compliance score of the operation result graph based on the calibration requirement compliance score group includes: determining the source role and score of each calibration requirement compliance score in the calibration requirement compliance score group of the operation result graph, wherein the source role is a tester or a product manager; and calculating the calibration requirement compliance score of the operation result graph based on the source role and score of each calibration requirement compliance score in the following manner:

[0021]

[0022]

[0023] Among them, T is the calibration requirement compliance score of the operation result graph, is the average score of the testers on the calibration requirement compliance of this operation result graph, is the mean score of the product manager on the calibration requirement compliance of this operation result graph, m T The number of testers who calibrate the requirement compliance score for this operation result graph, m P The number of product managers who calibrated the requirements compliance score for this operation result graph, is the score of the i-th tester for the calibration requirement compliance of this operation result graph, The j-th product manager's calibration requirement compliance score for this operation result graph.

[0024] In combination with the seventh possible implementation method of the first aspect, in the eighth possible implementation method of the first aspect, based on the expected result graph set and the operation result graph set, the requirement implementation compliance of the software product of the project to be evaluated is generated, including: for each test case in the completed test case set: determining the corresponding expected result graph from the expected result graph set, and determining the corresponding operation result graph from the operation result graph set; based on the expected result graph and the operation result graph, determining the test requirement compliance score of this test case; based on the test requirement compliance score corresponding to each test case, generating the requirement implementation compliance of the software product of the project to be evaluated.

[0025] In combination with the eighth possible implementation method of the first aspect, in the ninth possible implementation method of the first aspect, based on the test requirement compliance score corresponding to each test case, the requirement implementation compliance of the software product of the project to be evaluated is generated, including: performing a threshold judgment on the test requirement compliance score corresponding to each test case: if the test requirement compliance score corresponding to this test case reaches the threshold, determining that the function corresponding to the requirement corresponding to this test case in the software product is qualified; if the test requirement compliance score corresponding to this test case does not reach the threshold, determining that the function corresponding to the requirement corresponding to this test case in the software product is unqualified; calculating the pass rate of the functions corresponding to the requirements of all test cases in the software product as the requirement implementation compliance of the software product of the project to be evaluated, and marking all unqualified functions in the software product.

[0026] Beneficial effects:

[0027] 1. This solution obtains the product requirement specification, product page UI design, product page interaction design, a priori test case set (which can be an empty set or a non-empty set, and if it is an empty set, it is deemed that this priori test case set is not required) and software product of the project to be evaluated, and inputs the product requirement specification, product page UI design, product page interaction design and a priori test case set into the trained first model. The first model outputs a complete test case set, an operation instruction set and an expected result graph set, wherein the complete test case set includes the priori test case set, and each test case in the complete test case set corresponds to an operation instruction chain in the operation instruction set and an expected result graph in the expected result graph set, and each operation instruction chain contains at least one operation instruction; the complete test case set, operation instruction set and expected result graph set are input into the second model and the software product is called. The second model calls the corresponding operation instruction in the operation instruction set based on the complete test case set, runs the software product and automatically takes screenshots to form an operation result graph set, and generates the requirement implementation compliance of the software product of the project to be evaluated based on the expected result graph set and the operation result graph set. This approach can utilize the first model to generate a complete test case set based on the product requirement specification, product page UI design, product page interaction design, prior test case set, etc. (the complete test case set covers every requirement in the product requirement specification, which not only frees up the acceptance personnel to write test cases during acceptance, but also improves the acceptance coverage and effectively reduces omissions), operation instruction set and expected result atlas, and utilize the second model to call the software product for automatic testing, screenshots, and requirement implementation compliance assessment, etc., as a powerful auxiliary tool for testers and product managers to test and accept software products, and achieve efficient acceptance in a low-cost manner.

[0028] 2. The offline large model ChatGLM-6B is used for privatization training to obtain the first model and the second model. By collecting historical project data sets (including multiple historical project data, each historical project data includes software products, product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset, and expected result graph subset), the historical project data sets are divided into training sets and verification sets; the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset, and expected result graph subset of the historical project data in the training set are imported into the offline large model ChatGLM-6B for training, and the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset, and expected result graph subset of the historical project data in the verification set are imported into the trained offline large model ChatGLM-6B for verification and optimization, and the optimized offline large model ChatGLM-6B is obtained as the first model. The first model is obtained through the associated training of several collaborative sub-models, which allows each sub-model to perform relatively specialized work content, avoiding the problem of difficulty in refinement when a single model performs large-span work (the overall model usually considers overall efficiency, and the generalization effect is poor when it comes to work content with a large span), thereby improving the accuracy of the first model and improving the accuracy and reliability of the generated content.

[0029] 3. During the training of the first and second models, the generated content is evaluated to achieve feedback and updates on the model. This solution takes into account the perspectives of various roles in the acceptance process (testers and product managers have different perspectives and consider different factors during acceptance. The perspective and position of testers are the realization of the demand scenario and the possible risks involved in the demand, such as exception testing, functional usability testing, stability testing, etc.; while the perspective and position of product managers are more concerned with whether the designed software product requirements are realized and whether they meet expectations). During the feedback and update process, a solution of multi-role and multi-person scoring feedback during the training process is adopted to evaluate the generated content and provide feedback to the model to achieve verification and tuning of the model. When giving feedback, the differences in the perspectives of multiple roles and multiple people when scoring are taken into account. Different evaluation schemes are designed for different generated content to provide feedback, which is more in line with reality so that the model can be more accurate when conducting the final product demand implementation compliance assessment.

[0030] 4. For generated test cases (i.e., generative test cases), the emphasis is on testing, so the tester's scoring weight should be higher than that of the product manager. However, the variability of scores across different individuals must be considered. Therefore, when calculating the accuracy score for generative test cases, the consistency of scores across roles is also taken into account, dynamically assigning scoring weights to provide better feedback to the model and improve the accuracy of model-generated content. For generated operation instructions (generative operation instructions), the weights of testers and product managers tend to be balanced, so a mean calculation approach is adopted. However, for the evaluation of operation result graphs (which, in addition to requirement fulfillment, also focus on the degree of compliance with expectations after requirement fulfillment), the product manager's weight should be higher than that of testers. Therefore, when calculating the calibrated requirement conformance score for the operation result graph, the weight tends to be higher for the product manager. At the same time, the consistency of scores across roles is also taken into account, dynamically assigning scoring weights to provide better feedback to the model and improve the accuracy of the model's requirement fulfillment conformance assessment. Therefore, the model can be used as an auxiliary tool during acceptance, and artificial intelligence solutions can be used to effectively reduce the workload of personnel. It can usually ensure acceptance coverage and effectively reduce omissions. Based on the implementation compliance (test requirement compliance score) given by the model for each requirement, unqualified parts can be marked to facilitate key inspections by testers and product managers, and can also improve the reliability of acceptance.

[0031] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 A schematic diagram of the combination of the first model and the second model provided in an embodiment of the present application.

[0034] Figure 2 Flowchart for training the first model.

[0035] Figure 3 Flowchart for training the second model.

[0036] Figure 4 A flowchart of the artificial intelligence-based software product requirements implementation compliance assessment method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0038] Since the operation of this solution mainly depends on the operation of the first model and the second model, the first model and the second model are introduced here first.

[0039] See also Figure 1 , Figure 1 Schematic diagram of the combination of the first model and the second model provided in the embodiment of the present application. The first model is an integrated model (modelbase) that combines four sub-models (sub-model modelA, sub-model modelB, sub-model modelC, and sub-model modelD), while the second model relies on the content generated by the first model as input and ultimately outputs the compliance of the software product with the requirements.

[0040] To facilitate understanding of this solution, the training process of the first model and the second model is introduced here.

[0041] See also Figure 2 The training process of the first model may include step S11, step S12, and step S13.

[0042] First, step S11 may be executed.

[0043] Step S11: Obtain a historical project data set, where the historical project data set contains multiple historical project data. Each historical project data set contains a software product, a product requirement specification, a product page UI design, a product page interaction design, a test case subset, an operation instruction subset, and an expected result graph subset. The test case subset contains N test cases, the operation instruction subset contains N operation instruction chains, each operation instruction chain contains at least one operation instruction, and the expected result graph subset contains N expected result graphs.

[0044] In this embodiment, historical project data sets can be collected, and each historical project data set includes a software product (i.e., a software product that has been developed and tested and released, waiting for acceptance by the product manager), a product requirement specification (describing the various functions and various requirements that the product needs to implement), a product page UI design (i.e., a page prototype UI, with various elements. For example, various operation areas, display parts, etc.), a product page interaction design (describing the interaction of the product page, such as operation methods, page jump methods, etc.), a subset of test cases (test cases written by testers when conducting acceptance tests), a subset of operation instructions (operation instructions written according to test cases), and a subset of expected result diagrams (expected results determined based on the product page UI design and operation instructions, that is, the interface diagram expected to be achieved after performing an operation corresponding to a certain requirement in the product page UI design).

[0045] For example, this solution can collect a total of 100 historical project data from multiple provinces and multiple versions of this unit in the past. Each historical project data includes "Product Page UI Design", "Product Page Interaction Design", "Product Requirements Specification", "Test Case", as well as operation instructions and expected result diagrams during the acceptance process (of course, operation result diagrams can also be collected, but this solution adopts a manual calibration feedback tuning solution during the training process, so there is no need to collect operation result diagrams from past acceptances).

[0046] After the historical project data set is collected, step S12 may be executed.

[0047] Step S12: Divide the historical project dataset into a training set and a validation set.

[0048] In this embodiment, the historical project data set may be divided into a training set and a validation set in a ratio of 7:3.

[0049] Afterwards, step S13 may be executed.

[0050] Step S13: Import the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the training set into the offline large model ChatGLM-6B for training; import the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the verification set into the trained offline large model ChatGLM-6B for verification and optimization, and obtain the optimized offline large model ChatGLM-6B as the first model.

[0051] In this embodiment, the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the training set can be imported into the offline large model ChatGLM-6B for training.

[0052] For example, the product page UI design of historical project data in the training set is imported into the offline large model ChatGLM-6B (a generative large model framework open sourced by Tsinghua University) for private training, which is used to analyze various elements in the product page UI design (such as various operation areas, various display areas, etc.) to obtain sub-model modelA. The main function of sub-model modelA is to analyze and understand the UI design page.

[0053] The subset of product page interaction designs and operation instructions of historical project data in the training set is imported into the offline large model ChatGLM-6B for private training. This is used to analyze the correspondence between operation instructions and interactions in product page interaction designs, and obtain sub-model modelB.

[0054] The product requirement specifications of historical project data in the training set are imported into the offline large model ChatGLM-6B for private training to analyze product requirements and obtain the sub-model modelC.

[0055] In addition, a subset of test cases from historical project data in the training set is imported into the offline large model ChatGLM-6B for private training, which is used to analyze the test cases and obtain the sub-model modelD.

[0056] The four sub-models (sub-model modelA, sub-model modelB, sub-model modelC, sub-model modelD) each have their own focus and are potentially related. In order to make the best use of the four sub-models and complete more comprehensive tasks, it is necessary to combine the four sub-models and train an integrated model to play a central control and scheduling role to complete the given tasks in this plan.

[0057] For example, an association can be established between the offline large model ChatGLM-6B and sub-model modelA, sub-model modelB, sub-model modelC and sub-model modelD, and the product requirement specification, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the training set can be imported for association training to obtain the integrated model modelbase. The integrated model modelbase generates the generative test case corresponding to each requirement in the product requirement specification of the historical project data, the generative operation instruction chain corresponding to each generative test case and the generative expected result graph, and based on the test case subset, operation instruction subset and expected result graph subset of the historical project data, each generative test case of the historical project data, the generative operation instruction chain and the generative expected result graph corresponding to each generative test case are fed back and updated.

[0058] Sub-model modelA analyzes the product page UI design and learns the various elements of each page within it. Sub-model modelB analyzes the correspondence between operational instructions and interactions within the product page interaction design. Sub-model modelC analyzes product requirements, and sub-model modelD analyzes test cases. Under the overall coordination of the integrated model modelbase, sub-models modelA, modelB, modelC, and modelD collaborate to analyze product requirements based on the product requirements specification. They analyze the relationship between requirements and test cases based on a subset of product requirements and test cases. They analyze the relationship between test cases and operational instructions based on a subset of test cases and the operational instruction subset, combined with the product page interaction design. Finally, they analyze the relationship between operational instructions and the expected result graph based on the subset of operational instructions and the subset of expected result graphs, combined with the product page UI design. This allows analysis of the product requirements specification, product page UI design, product page interaction design, a subset of test cases, a subset of operational instructions, and a subset of expected result graphs within historical project data, significantly improving the reliability and accuracy of generated content.

[0059] When importing the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the verification set into the trained offline large model ChatGLM-6B for verification and optimization, the product requirement specifications, product page UI design, and product page interaction design of the historical project data in the verification set can be imported into the integrated model modelbase.

[0060] Then, for each historical project data in the validation set:

[0061] The integrated model modelbase can generate generative test cases corresponding to each requirement in the product requirement specification, generative operation instruction chains corresponding to each generative test case, and generative expected result diagrams based on the product requirement specification, product page UI design, and product page interaction design of historical project data.

[0062] For the generative test cases corresponding to each requirement: the accuracy score group of the generative test cases can be obtained, and the accuracy score of this generative test case can be determined based on the accuracy score group. Then, based on the accuracy score of this generative test case and the test case corresponding to this requirement in the test case set, the integrated model modelbase is verified and optimized. Among them, the accuracy score group includes the accuracy scores of multiple testers and / or product managers for this generative test case.

[0063] Specifically, because testers and product managers have different perspectives and consider different factors during acceptance, the testers' perspective and position are the realization of the demand scenario and the possible risks involved in the demand, such as exception testing, functional usability testing, and stability testing; while the product manager's perspective and position are more concerned with whether the designed software product requirements are realized and whether they meet expectations. Then, for generative test cases, the source role and score of each accuracy score in the accuracy score group of the generative test case can be determined, where the source role is the tester or product manager. Since this scenario is an accuracy assessment of the test case, the weight of the tester should be higher than that of the product manager. At the same time, considering the impact of score consistency (the higher the consistency, the stronger the consensus of people in this role on the result, and it is usually more credible), based on the source role and score of each accuracy score, the accuracy score of this generative test case is calculated in the following way:

[0064]

[0065] Among them, R is the accuracy score of the generated test case, is the mean accuracy score of the testers for this generated test case, is the mean accuracy score of the product manager for this generative test case, n T The number of testers who rated the accuracy of this generated test case, n P The number of product managers who rated the accuracy of this generative test case, is the accuracy score of the generative test case given by the i-th tester, The accuracy score of the generative test case given by the j-th product manager.

[0066] For the generative operation instruction chain corresponding to each generative test case: obtain the accuracy score group of the generative operation instruction chain, and determine the accuracy score of this generative operation instruction chain based on the accuracy score group, and then verify and optimize the integrated model modelbase based on the accuracy score of this generative operation instruction chain and the operation instruction chain of the corresponding test case, where the accuracy score group includes the accuracy scores of multiple testers and / or product managers for this generative operation instruction chain.

[0067] For operation instructions, the weights of testers and product managers should be balanced. Therefore, the accuracy score of the generated operation instruction chain can be calculated based on the score of each accuracy score in the accuracy score group of the generated operation instruction chain in the following way:

[0068]

[0069] Among them, S is the accuracy score of the generated operation instruction chain, n is the total number of people who score the accuracy of this generated operation instruction chain, and the people are testers or product managers. i The accuracy score of the generated operation instruction chain for the i-th person.

[0070] For each generative test case's corresponding generative expected result graph, the integrated model modelbase can be verified and optimized based on the generative expected result graph and the corresponding expected result graph in the expected result graph subset. Since the generative expected result graph originates from the UI design page in the product page UI design, direct feedback can be used to verify the integrated model modelbase (depending on the operation instructions). This yields the optimized offline large model ChatGLM-6B (an integrated model modelbase that combines four sub-models) as the first model.

[0071] On this basis, it is necessary to train a second model to evaluate the compliance of product requirements. The second model is also obtained by private training using the offline large model ChatGLM-6B. Figure 3 The training process of the second model may include step S21 and step S22.

[0072] First, step S21 may be executed.

[0073] Step S21: Generate a corresponding software call path based on the software product of each historical project data in the historical project data set, and import the software call path into the offline large model ChatGLM-6B.

[0074] In this embodiment, for each software product of historical project data, a corresponding software call path can be generated and imported into the offline large model ChatGLM-6B for the offline large model ChatGLM-6B to call to run the corresponding software product.

[0075] Afterwards, step S22 may be executed.

[0076] Step S22: For each historical project data, the test case subset, operation instruction subset and expected result graph subset of the historical project data are imported into the offline large model ChatGLM-6B, and the software call path is called by the offline large model ChatGLM-6B to run the software product. For each test case in the test case subset, the software product is controlled to execute the operation instruction chain of the test case, and a screenshot of the interface of the software product after executing the operation instruction chain is taken to obtain the operation result graph corresponding to the test case. Based on the operation result graph and the corresponding expected result graph, the test requirement compliance score of this operation result graph is determined, and the calibration requirement compliance score group of this operation result graph is obtained, and the calibration requirement compliance score of this operation result graph is determined based on the calibration requirement compliance score group. Based on the calibration requirement compliance score and the test requirement compliance score of this operation result graph, the offline large model ChatGLM-6B is fed back and updated, and finally the second model that has been trained is obtained.

[0077] In this embodiment, for each historical project data, a subset of test cases, a subset of operation instructions, and a subset of expected result diagrams of the historical project data can be imported into the offline large model ChatGLM-6B, and the software call path is called by the offline large model ChatGLM-6B to run the software product.

[0078] Then, for each test case in the test case set, the software product can be controlled to execute the operation instruction chain of the test case, and a screenshot of the software product interface after executing the operation instruction chain can be taken to obtain the operation result diagram corresponding to the test case.

[0079] Then, based on the operation result graph and the corresponding expected result graph, the test requirement compliance score of this operation result graph is determined (this can rely on the offline large model ChatGLM-6B to evaluate the similarity between the two). In addition, the calibration requirement compliance score group of this operation result graph is obtained, and the calibration requirement compliance score of this operation result graph is determined based on the calibration requirement compliance score group. Based on the calibration requirement compliance score and test requirement compliance score of this operation result graph, the offline large model ChatGLM-6B is fed back and updated, and finally the second model is trained.

[0080] Specifically, the source role and score of each calibration requirement compliance score in the calibration requirement compliance score group of the operation result graph can be determined, where the source role is the tester or product manager. Since this scenario is a requirement compliance assessment, the weight of the product manager should be higher than that of the tester. At the same time, considering the impact of score consistency (the higher the consistency, the stronger the consensus of people in this role on the result, and generally more credible), the calibration requirement compliance score of this operation result graph can be calculated based on the source role and score of each calibration requirement compliance score in the following way:

[0081]

[0082] Among them, T is the calibration requirement compliance score of the operation result graph, is the average score of the testers on the calibration requirement compliance of this operation result graph, is the mean score of the product manager on the calibration requirement compliance of this operation result graph, m T The number of testers who calibrate the requirement compliance score for this operation result graph, m P The number of product managers who calibrated the requirements compliance score for this operation result graph, is the score of the i-th tester for the calibration requirement compliance of this operation result graph, The j-th product manager's calibration requirement compliance score for this operation result graph.

[0083] After completing the training of the first model and the second model, the connection between the first model and the second model can be established so that the first model outputs its generated test case set, operation instruction set and expected result atlas (because there is no high-coverage test case set, operation instruction set and expected result atlas, etc. during application, it is necessary to rely on the trained first model to automatically generate. The names of the generated content during training and during application are different, mainly for the purpose of distinction, and should not be regarded as a limitation of this application) to the second model, and the second model evaluates the compliance of product requirements based on these test case sets, operation instruction sets and expected result atlases. Of course, after the first model and the second model are combined, they can be applied directly, or they can be verified with a part of the historical project data set (or a small amount of new historical project data can be collected for verification) to evaluate the effectiveness of the model.

[0084] Based on this, an artificial intelligence-based software product requirement implementation compliance assessment method can be run to realize intelligent software product requirement implementation compliance assessment.

[0085] See also Figure 4 , Figure 4Flowchart of a method for evaluating the compliance of software product requirements based on artificial intelligence. In this embodiment, the method for evaluating the compliance of software product requirements based on artificial intelligence may include steps S31, S32, and S33.

[0086] First, step S31 may be executed.

[0087] Step S31: Obtain the product requirement specification, product page UI design, product page interaction design, a priori test case set and software product of the project to be evaluated, wherein the a priori test case set is an empty set or a non-empty set, and the test cases in the a priori test case set are generated by the product manager and / or testers to test some functions of this software product.

[0088] In this embodiment, the product requirement specification, product page UI design, product page interaction design, a priori test case set and software product (software that has been developed and tested and released) of the project to be evaluated can be obtained. The a priori test case set is an empty set or a non-empty set. Generally speaking, there are some test cases in the testing process, so this a priori test case set can be formed, but the acceptance coverage is not enough. Of course, this a priori test case set can also be not sorted as input (that is, the a priori test case set is an empty set). When the a priori test case set is a non-empty set, the test cases are generated by the product manager and / or tester to test some functions of this software product.

[0089] After obtaining the product requirement specification, product page UI design, product page interaction design, prior test case set and software product, step S32 can be executed.

[0090] Step S32: Input the product requirement specification, product page UI design, product page interaction design and prior test case set into the trained first model, and output the completed test case set, operation instruction set and expected result graph set through the first model, wherein the completed test case set includes the prior test case set, and each test case in the completed test case set corresponds to an operation instruction chain in the operation instruction set and an expected result graph in the expected result graph set, and each operation instruction chain contains at least one operation instruction.

[0091] In this embodiment, the product requirement specification, product page UI design, product page interaction design and a priori test case set (this item is not input when the priori test case set is an empty set) can be input into the trained first model, and the first model is used to output the completed test case set (i.e., the generated test case set. If the priori test case set is a non-empty set, then the completed test case set includes all the test cases in the priori test case set), the operation instruction set and the expected result graph set. Each test case in the completed test case set corresponds to an operation instruction chain in the operation instruction set and an expected result graph in the expected result graph set. Each operation instruction chain contains at least one operation instruction (because a test case generally contains multiple operation steps, which correspond to a series of operation instructions. This solution is named an operation instruction chain, which is actually the operation instruction corresponding to the interactive operation required to implement this test case).

[0092] After obtaining the completed test case set, operation instruction set, and expected result graph set output by the first model, step S33 may be executed.

[0093] Step S33: Input the completed test case set, operation instruction set and expected result atlas into the second model and call the software product. The second model calls the corresponding operation instructions in the operation instruction set based on the completed test case set to run the software product and automatically take screenshots to form an operation result atlas. Based on the expected result atlas and the operation result atlas, the requirement implementation compliance of the software product of the project to be evaluated is generated.

[0094] In this embodiment, the completed test case set, the operation instruction set, and the expected result graph set can be input into the second model and the software product can be called. The second model then calls the corresponding operation instructions in the operation instruction set based on the completed test case set to run the software product and automatically takes screenshots, forming an operation result graph set. Based on the expected result graph set and the operation result graph set, the requirements implementation compliance of the software product of the project to be evaluated is generated. The software product calling process can be described in the previous section, that is, the software call path of the software product is generated and imported into the second model for the second model to call. The software product is run using the operation instructions and screenshots are taken to obtain the operation result graph, forming the operation result graph set.

[0095] Afterwards, the requirements implementation compliance of the software product of the project to be evaluated can be generated based on the expected result atlas and the operation result atlas.

[0096] Exemplarily, for each test case in the completed test case set: the corresponding expected result graph can be determined from the expected result graph set, and the corresponding operation result graph can be determined from the operation result graph set. Then, based on the expected result graph and the operation result graph, the test requirement compliance score of this test case is determined (i.e., the test requirement compliance score generated as introduced in the training process above is updated through feedback of the calibration requirement compliance score to improve the accuracy of the model evaluation). Afterwards, based on the test requirement compliance score corresponding to each test case, the requirement implementation compliance of the software product of the project to be evaluated can be generated.

[0097] Specifically, a threshold value can be used to determine the test requirement compliance score for each test case:

[0098] If the test requirement compliance score corresponding to this test case reaches a threshold (for example, 90%, 90% here is just an example, and different scoring systems correspond to different values. For example, if a 10-point system is used, the threshold is 9 points), it is determined that the function corresponding to the requirement of this test case in the software product is qualified.

[0099] If the test requirement compliance score corresponding to this test case does not reach the threshold, it is determined that the function corresponding to the requirement corresponding to this test case in the software product is unqualified.

[0100] Then, the pass rate of the functions corresponding to the requirements of all test cases in the software product can be calculated as the compliance degree of the requirements of the software product of the project to be evaluated, and all unqualified functions in the software product can be marked so that acceptance personnel (product managers, testers, etc.) can focus on and test them.

[0101] In summary, the embodiment of the present application provides a method for evaluating the compliance of software product requirements based on artificial intelligence, which obtains the product requirement specification, product page UI design, product page interaction design, a priori test case set (which can be an empty set or a non-empty set, and when it is an empty set, it is deemed that this priori test case set is not needed) and the software product of the project to be evaluated, and inputs the product requirement specification, product page UI design, product page interaction design and a priori test case set into the trained first model, and outputs the completed test case set, operation instruction set and expected result graph set through the first model, wherein the completed test case The set includes a priori test case set, and each test case in the completion test case set corresponds to an operation instruction chain in the operation instruction set and an expected result diagram in the expected result diagram set, and each operation instruction chain contains at least one operation instruction; the completion test case set, the operation instruction set and the expected result diagram set are input into the second model and the software product is called, and the second model calls the corresponding operation instruction in the operation instruction set based on the completion test case set to run the software product and automatically take a screenshot to form an operation result diagram set, and based on the expected result diagram set and the operation result diagram set, the requirement implementation compliance of the software product of the project to be evaluated is generated. This approach can utilize the first model to generate a complete test case set based on the product requirement specification, product page UI design, product page interaction design, prior test case set, etc. (the complete test case set covers every requirement in the product requirement specification, which not only frees up the acceptance personnel to write test cases during acceptance, but also improves the acceptance coverage and effectively reduces omissions), operation instruction set and expected result atlas, and utilize the second model to call the software product for automatic testing, screenshots, and requirement implementation compliance assessment, etc., as a powerful auxiliary tool for testers and product managers to test and accept software products, and achieve efficient acceptance in a low-cost manner.

[0102] The offline large model ChatGLM-6B is used for privatization training to obtain the first model and the second model. By collecting historical project data sets (including multiple historical project data, each historical project data includes software products, product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset, expected result diagram subset and software products), the historical project data sets are divided into training sets and verification sets; the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result diagram subset of the historical project data in the training set are imported into the offline large model ChatGLM-6B for training, and the product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result diagram subset of the historical project data in the verification set are imported into the trained offline large model ChatGLM-6B for verification and optimization, and the optimized offline large model ChatGLM-6B is obtained as the first model. The first model is obtained through the associated training of several collaborative sub-models, which allows each sub-model to perform relatively specialized work content, avoiding the problem of difficulty in refinement when a single model performs large-span work (the overall model usually considers overall efficiency, and the generalization effect is poor when it comes to work content with a large span), thereby improving the accuracy of the first model and improving the accuracy and reliability of the generated content.

[0103] During the training of the first and second models, the generated content is evaluated to provide feedback and updates to the models. This solution takes into account the perspectives of various roles in the acceptance process (testers and product managers have different perspectives and consider different factors during acceptance. The tester's perspective and position are the realization of the demand scenario and the possible risks involved in the requirements, such as exception testing, functional usability testing, and stability testing; while the product manager's perspective and position are more concerned with whether the designed software product requirements are met and meet expectations). During the feedback and update process, a multi-role and multi-person scoring and feedback solution is used during the training process to evaluate the generated content and provide feedback to the model to achieve model verification and tuning. When providing feedback, the differences in the perspectives of multiple roles and multiple people are taken into account. Different evaluation solutions are designed for different generated content to provide feedback, which is more realistic and allows the model to be more accurate when conducting the final product requirement implementation compliance assessment.

[0104] For generated test cases (i.e., generative test cases), the emphasis is on testing, so the tester's score weight should be higher than the product manager's. However, the variability of scores across different individuals must be considered. Therefore, when calculating the accuracy score for generative test cases, the consistency of scores across roles is also considered, dynamically assigning scoring weights to provide better feedback to the model and improve the accuracy of model-generated content. For generated operation instructions (generative operation instructions), the weights assigned to testers and product managers tend to be balanced, so a mean calculation approach is adopted. However, for the evaluation of operation result graphs (which, in addition to requirement fulfillment, also focus on the degree of compliance with expectations after requirement fulfillment), the product manager's weight should be higher than that of the tester. Therefore, when calculating the calibrated requirement conformance score for the operation result graph, a higher weight is assigned to the product manager. Furthermore, the consistency of scores across roles is also considered, dynamically assigning scoring weights to provide better feedback to the model and improve the accuracy of the model's requirement fulfillment conformance assessment. Therefore, the model can be used as an auxiliary tool during acceptance, and artificial intelligence solutions can be used to effectively reduce the workload of personnel. It can usually ensure acceptance coverage and effectively reduce omissions. Based on the implementation compliance (test requirement compliance score) given by the model for each requirement, unqualified parts can be marked to facilitate key inspections by testers and product managers, and can also improve the reliability of acceptance.

[0105] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0106] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Persons skilled in the art will readily appreciate that various modifications and variations to the present application are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A software product requirement implementation compliance assessment method based on artificial intelligence, characterized in that: include: Obtain the product requirements specification, product page UI design, product page interaction design, prior test case set, and software product of the project to be evaluated. The prior test case set may be an empty set or a non-empty set, and the test cases in the prior test case set are generated by the product manager and / or testers to test some functions of the software product. Input the product requirement specification, product page UI design, product page interaction design, and a priori test case set into the trained first model, and output a completion test case set, an operation instruction set, and an expected result graph set through the first model, wherein the completion test case set includes the priori test case set, and each test case in the completion test case set corresponds to an operation instruction chain in the operation instruction set and an expected result graph in the expected result graph set, and each operation instruction chain includes at least one operation instruction; The completed test case set, operation instruction set and expected result atlas are input into the second model and the software product is called. The second model calls the corresponding operation instructions in the operation instruction set based on the completed test case set to run the software product and automatically take screenshots to form an operation result atlas. Based on the expected result atlas and the operation result atlas, the requirement implementation compliance of the software product of the project to be evaluated is generated.

2. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 1, characterized in that: The first model is obtained by private training using the offline large model ChatGLM-6B. The training process of the first model is as follows: Obtain a historical project dataset, where the historical project dataset contains multiple historical project data. Each historical project data set includes a software product, a product requirement specification, a product page UI design, a product page interaction design, a test case subset, an operation instruction subset, and an expected result graph subset. The test case subset contains N test cases, the operation instruction subset contains N operation instruction chains, each operation instruction chain contains at least one operation instruction, and the expected result graph subset contains N expected result graphs. Divide the historical project dataset into training set and validation set; The product requirement specifications, product page UI design, product page interaction design, a subset of test cases, a subset of operation instructions, and a subset of expected result diagrams of the historical project data in the training set are imported into the offline large model ChatGLM-6B for training. The product requirement specifications, product page UI design, product page interaction design, a subset of test cases, a subset of operation instructions, and a subset of expected result diagrams of the historical project data in the verification set are imported into the trained offline large model ChatGLM-6B for verification and optimization, and the optimized offline large model ChatGLM-6B is obtained as the first model.

3. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 2, characterized in that: The product requirement specifications, product page UI design, product page interaction design, test case subset, operation instruction subset, and expected result graph subset of historical project data in the training set are imported into the offline large model ChatGLM-6B for training, including: Import the product page UI design of historical project data in the training set into the offline large model ChatGLM-6B for private training. This is used to analyze various elements in the product page UI design and obtain sub-model modelA. Import the product page interaction design and operation instruction subset of historical project data in the training set into the offline large model ChatGLM-6B for private training. This is used to analyze the correspondence between operation instructions and interactions in the product page interaction design, and obtain sub-model modelB. Import the product requirement specifications of historical project data in the training set into the offline large model ChatGLM-6B for private training to analyze product requirements and obtain sub-model modelC; Import the test case subset of historical project data in the training set into the offline large model ChatGLM-6B for private training to analyze the test cases and obtain the sub-model modelD; Establish an association between the offline large model ChatGLM-6B and sub-model modelA, sub-model modelB, sub-model modelC and sub-model modelD, and import the product requirement specification, product page UI design, product page interaction design, test case subset, operation instruction subset and expected result graph subset of the historical project data in the training set for association training to obtain the integrated model modelbase. Through the integrated model modelbase, generate the generative test case corresponding to each requirement in the product requirement specification of the historical project data, the generative operation instruction chain corresponding to each generative test case and the generative expected result graph, and based on the test case subset, operation instruction subset and expected result graph subset of the historical project data, feedback update is performed on each generative test case of the historical project data, the generative operation instruction chain corresponding to each generative test case and the generative expected result graph.

4. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 3, characterized in that: Import the product requirements specifications, product page UI design, product page interaction design, test case subset, operation instruction subset, and expected result graph subset of historical project data from the validation set into the trained offline large model ChatGLM-6B for validation and optimization, including: Import the product requirement specifications, product page UI design, and product page interaction design of the verified historical project data into the integrated model modelbase; For each historical project data in the validation set, the model base is integrated with the product requirements specification, product page UI design, and product page interaction design based on the historical project data to generate generative test cases corresponding to each requirement in the product requirements specification, generative operation instruction chains corresponding to each generative test case, and generative expected result diagrams. For each generative test case corresponding to a requirement: obtain the accuracy score group of the generative test case, and determine the accuracy score of this generative test case based on the accuracy score group. Then, based on the accuracy score of this generative test case and the test case corresponding to this requirement in the test case collection, verify and optimize the integrated model modelbase. The accuracy score group includes the accuracy scores of multiple testers and / or product managers for this generative test case. For each generative operation instruction chain corresponding to a generative test case: obtain the accuracy score group of the generative operation instruction chain, and determine the accuracy score of this generative operation instruction chain based on the accuracy score group. Then, based on the accuracy score of this generative operation instruction chain and the operation instruction chain of the corresponding test case, verify and optimize the integrated model modelbase, where the accuracy score group includes the accuracy scores of multiple testers and / or product managers for this generative operation instruction chain; For the generative expected result graph corresponding to each generative test case: based on the generative expected result graph and the corresponding expected result graph in the expected result graph subset, the integrated model modelbase is verified and optimized.

5. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 4, characterized in that: The accuracy score of this generated test case is determined based on the accuracy score group, including: Determine the source role and score of each accuracy score in the accuracy score group of the generative test case, where the source role is the tester or product manager; Based on the source role and score for each accuracy score, the accuracy score for this generative test case is calculated as follows: Among them, R is the accuracy score of the generated test case, is the mean accuracy score of the testers for this generated test case, is the mean accuracy score of the product manager for this generative test case, n T The number of testers who rated the accuracy of this generated test case, n P The number of product managers who rated the accuracy of this generative test case, is the accuracy score of the generative test case given by the i-th tester, The accuracy score of the generative test case given by the j-th product manager.

6. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 4, characterized in that: The accuracy score of the generative operation instruction chain is determined based on the accuracy score group, including: Based on the accuracy score of each item in the accuracy score group of the generated operation instruction chain, the accuracy score of this generated operation instruction chain is calculated in the following way: Among them, S is the accuracy score of the generated operation instruction chain, n is the total number of people who score the accuracy of this generated operation instruction chain, and the people are testers or product managers. i The accuracy score of the generated operation instruction chain for the i-th person.

7. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 2, characterized in that: The second model is obtained by private training using the offline large model ChatGLM-6B. The training process of the second model is as follows: Generate the corresponding software call path based on the software product of each historical project data in the historical project dataset, and import the software call path into the offline large model ChatGLM-6B; For each historical project data, the test case subset, operation instruction subset and expected result graph subset of the historical project data are imported into the offline large model ChatGLM-6B, and the software call path is called by the offline large model ChatGLM-6B to run the software product. For each test case in the test case subset, the software product is controlled to execute the operation instruction chain of the test case, and a screenshot of the interface of the software product after executing the operation instruction chain is taken to obtain the operation result graph corresponding to the test case. Based on the operation result graph and the corresponding expected result graph, the test requirement compliance score of this operation result graph is determined, and the calibration requirement compliance score group of this operation result graph is obtained, and the calibration requirement compliance score of this operation result graph is determined based on the calibration requirement compliance score group. Based on the calibration requirement compliance score and the test requirement compliance score of this operation result graph, the offline large model ChatGLM-6B is fed back and updated, and finally the second model that has been trained is obtained.

8. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 7, characterized in that: The calibration requirement compliance score of the operation result graph is determined based on the calibration requirement compliance score group, including: Determine the source role and score of each calibration requirement compliance score in the calibration requirement compliance score group of the operation result graph, where the source role is a tester or a product manager; Based on the source role and score of each calibration requirement compliance score, the calibration requirement compliance score of this operation result graph is calculated in the following way: Among them, T is the calibration requirement compliance score of the operation result graph, is the average score of the testers on the calibration requirement compliance of this operation result graph, is the mean score of the product manager on the calibration requirement compliance of this operation result graph, m T The number of testers who calibrate the requirement compliance score for this operation result graph, m P The number of product managers who calibrated the requirements compliance score for this operation result graph, is the score of the i-th tester for the calibration requirement compliance of this operation result graph, The j-th product manager's calibration requirement compliance score for this operation result graph.

9. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 8, characterized in that: Based on the expected result atlas and the operational result atlas, the requirements implementation compliance of the software product of the project to be evaluated is generated, including: For each test case in the completed test case set: Determining a corresponding expected result graph from the expected result graph set, and determining a corresponding operation result graph from the operation result graph set; Based on the expected result diagram and the operation result diagram, determine the test requirement compliance score of this test case; Based on the test requirement compliance score corresponding to each test case, the requirement implementation compliance of the software product of the project to be evaluated is generated.

10. The method for evaluating the compliance of software product requirements based on artificial intelligence according to claim 9, characterized in that: Based on the test requirement compliance score corresponding to each test case, the software product requirement implementation compliance of the project to be evaluated is generated, including: Perform threshold judgment on the test requirement compliance score corresponding to each test case: If the test requirement compliance score of this test case reaches the threshold, it is determined that the function corresponding to the requirement of this test case in the software product is qualified; If the test requirement compliance score corresponding to this test case does not reach the threshold, it is determined that the function corresponding to the requirement of this test case in the software product is unqualified; Calculate the pass rate of the functions corresponding to the requirements of all test cases in the software product as the compliance degree of the software product requirements of the project to be evaluated, and mark all unqualified functions in the software product.

Citation Information

Patent Citations

  • Software automation testing method, system and device

    CN117421231A

  • Test case generation method and device, terminal equipment and readable storage medium

    CN117707922A