Question and answer software testing method, device and equipment based on large language model
By extracting information from existing data sets and generating high-quality test cases using large language models, the naturalness and accuracy problems in question-and-answer software testing are solved, and more efficient defect detection is achieved.
Patent Information
- Application Number
- CN202510259437.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-11
AI Technical Summary
The test cases generated by existing Q&A software testing methods are poor in nature, cannot effectively trigger defects, and have low accuracy in defect detection.
Using a large language model to extract entity and relationship information from the background text of an existing dataset, generate high-quality test cases, and detect defects through semantic similarity and background text consistency scores.
The generated test cases are more natural and can effectively trigger the defects of the Q&A software, improving the testing efficiency and accuracy.
Smart Images

Figure CN120295906A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of Q&A software, and particularly to a method, device, and equipment for testing Q&A software based on large language models. Background Art
[0002] Closed-world Q&A software is software that accepts both background text and natural language questions as input and answers the input questions based on the input text. This type of software has important applications in enterprise internal knowledge management, medical consultation, user support, etc. Therefore, detecting whether the Q&A software contains defects has become an important task. In previous work, developers needed to manually mark the results of each test case, which would incur huge costs. So how to automatically generate a large number of test cases has received extensive attention from researchers.
[0003] In current methods for automatically generating test cases, mainly the original dataset is used as a seed set to design a series of metamorphic relations for generating new test cases. QAQA designs a series of sentence-level metamorphic relations on questions and background text respectively based on the idea that adding irrelevant sentences does not affect the question answer, and generates a large number of new test cases. QATest designs a complete fuzz testing framework for Q&A software, designs mutation methods at multiple levels from words to sentences based on the idea of semantic invariance, and conducts quality evaluation from two perspectives of coverage and doubt degree, and obtains new generated test cases after screening.
[0004] Considering existing work, they mainly generate new test cases with consistent expected answers through modification methods based on the questions in existing datasets (such as BoolQ, Squad2, etc.), with a small amount of addition, deletion, or synonym replacement. The test cases obtained based on this type of metamorphic method cannot guarantee their naturalness, and the actual value of the detected defects is relatively low. Summary of the Invention
[0005] This application provides a method, device, and equipment for testing Q&A software based on large language models to solve the problems in the above background art.
[0006] In a first aspect, this application provides a method for testing Q&A software based on large language models, including:
[0007] Extract information from the background text of the existing dataset as the standard answer for generating questions;
[0008] Use the large language model to ask questions based on the standard answer to generate questions as test cases;
[0009] Run the test cases on the Q&A software to be tested to obtain the answers of the Q&A software to be tested.
[0010] Further, extracting information from the background text of the existing dataset includes:
[0011] Extracting entity information and relationship information from the background text of the existing dataset, where the entity information includes nouns, verbs, noun phrases, and verb phrases, and the relationship information includes relationship triples.
[0012] Further, extracting entity information and relationship information from the background text of the existing dataset includes:
[0013] Performing part-of-speech and constituent analysis on gerund phrases and gerunds, setting screening conditions based on the dependency relationship of the sentence for screening, and outputting the extracted grammatical information;
[0014] Combining multiple examples and the chain of thought to construct a prompt for relationship extraction, and based on the prompt, using a large language model to obtain initial relationship triples.
[0015] Further, using the large language model to ask questions based on the standard answer to generate questions as test cases includes:
[0016] Inputting the standard answer combined with the background text into the large language model, and constructing a prompt in the format of instruction + restriction + multiple examples + chain of thought to generate questions as test cases.
[0017] Further, after using the large language model to ask questions based on the standard answer to generate questions as test cases, it further includes:
[0018] Screening the questions generated by the large language model by designing rules and re-asking the large language model.
[0019] Further, after running the test cases on the question-and-answer software to be tested and obtaining the answers of the question-and-answer software to be tested, it further includes:
[0020] Comparing whether the standard answer is consistent with the answer of the question-and-answer software to be tested to detect whether the test cases have discovered defects in the question-and-answer software to be tested.
[0021] Further, comparing whether the standard answer is consistent with the answer of the question-and-answer software to be tested to detect whether the test cases have discovered defects in the question-and-answer software to be tested includes:
[0022] Calculating the semantic similarity between the standard answer and the answer of the question-and-answer software to be tested;
[0023] If the semantic similarity does not reach the first preset threshold, using the large language model to perform a consistency score on the two answers based on the background text;
[0024] If the score is lower than the second preset threshold, it is determined that the test case has discovered a defect in the question-and-answer software to be tested.
[0025] In a second aspect, the present application provides a question-and-answer software testing device based on a large language model, including:
[0026] An information extraction module, configured to extract information from the background text of an existing data set as the standard answer for generating questions;
[0027] A test case generation module, configured to use the large language model to ask questions based on the standard answer to generate questions as test cases;
[0028] A software testing module, configured to run the test case on the question-and-answer software to be tested to obtain the answer of the question-and-answer software to be tested.
[0029] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned question-and-answer software testing method based on a large language model is implemented.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned question-and-answer software testing method based on a large language model is implemented.
[0031] The above technical solutions of the present application have the following advantages:
[0032] The question-and-answer software testing method based on a large language model provided in the first aspect of the present application extracts information from the background text of an existing data set as the standard answer for generating questions, uses the large language model to ask questions based on the standard answer to generate questions as test cases, and runs the test case on the question-and-answer software to be tested to obtain the answer of the question-and-answer software to be tested, making the generated test cases more natural and capable of triggering defects more effectively, thereby improving the test efficiency.
[0033] It can be understood that the beneficial effects of the above second aspect, third aspect, and fourth aspect can refer to the relevant descriptions in the above first aspect and will not be elaborated here. Description of the Drawings
[0034] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is a flowchart of the question-answering software testing method based on a large language model provided by the present application;
[0036] Figure 2 It is a schematic diagram of the principle of the question-answering software testing method based on a large language model provided by the present application;
[0037] Figure 3 It is a structural diagram of the question-answering software testing device based on a large language model provided by the present application;
[0038] Figure 4 It is a schematic structural diagram of the electronic device provided by the present application. Specific Embodiments
[0039] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0040] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0041] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0042] References to "one embodiment" or "some embodiments" in the description of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc., which appear at different places in this specification, do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized. "Plurality" means "two or more".
[0043] The following describes in detail the specific implementation manners of this application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate this application, but are not used to limit the scope of this application.
[0044] As Figure 1 shown, the embodiment of this application provides a method for testing a question-and-answer software based on a large language model, which specifically includes the following steps: extracting information from the background text of an existing dataset as the standard answer for generating questions; using the large language model to ask questions based on the standard answer to generate questions as test cases; running the test cases on the question-and-answer software to be tested to obtain the answers of the question-and-answer software to be tested.
[0045] In some embodiments, the extracting information from the background text of an existing dataset includes: extracting entity information and relationship information from the background text of an existing dataset, where the entity information includes nouns, verbs, noun phrases, and verb phrases, and the relationship information includes relationship triples.
[0046] In some embodiments, the extracting entity information and relationship information from the background text of an existing dataset includes: performing part-of-speech and constituent analysis on gerund phrases and gerunds, setting screening conditions according to the dependency relationship of sentences for screening, and outputting the extracted syntactic information; constructing prompt words for relationship extraction by combining multi-instance and chain of thought, and based on the prompt words, using the large language model to obtain initial relationship triples.
[0047] In some embodiments, the using the large language model to ask questions based on the standard answer to generate questions as test cases includes: inputting the standard answer in combination with the background text into the large language model, and constructing prompt words in the format of instruction + restriction + multi-instance + chain of thought to generate questions as test cases.
[0048] In some embodiments, after using the large language model to ask questions based on the standard answer and generate questions as test cases, the method further includes: screening the questions generated by the large language model through design rules and re-asking the large language model.
[0049] In some embodiments, after running the test case on the question-and-answer software to be tested and obtaining the answer of the question-and-answer software to be tested, the method further includes: comparing whether the standard answer is consistent with the answer of the question-and-answer software to be tested to detect whether the test case discovers a defect of the question-and-answer software to be tested.
[0050] In some embodiments, comparing whether the standard answer is consistent with the answer of the question-and-answer software to be tested to detect whether the test case discovers a defect of the question-and-answer software to be tested includes: calculating the semantic similarity between the standard answer and the answer of the question-and-answer software to be tested; if the semantic similarity does not reach the first preset threshold, using the large language model to perform a consistency score on the two answers based on the background text; if the score is lower than the second preset threshold, determining that the test case discovers a defect of the question-and-answer software to be tested.
[0051] Existing work is mainly obtained by modifying the questions in the original dataset. However, based on the existing metamorphic method, it is impossible to ensure the naturalness of the generated questions. And the construction method of inserting irrelevant information in QAQA will generate a large number of questions with inconsistent semantics before and after. Considering that the purpose of designing test cases is to imitate people's daily usage, even if the questions generated by the existing methods can induce defects, their actual value is very limited. In addition, the modification of existing work is mainly based on questions, and the background articles in the dataset are not efficiently utilized. In actual use, users may ask questions about any part of the article, and the test cases of the existing methods cannot well cover the possible question range of users.
[0052] Since the test cases constructed by existing work are all semantically similar to the original questions in the dataset, based on the idea of metamorphic testing, if the answer of the software under test to the newly generated test case is inconsistent with the original answer, it is considered that a defect of the software under test is discovered. And the method of judging semantic consistency is to calculate semantic similarity using a pre-trained model or perform semantic judgment on the words in it. Such a detection method will produce misjudgment situations, mainly due to two reasons: one is limited by the performance of the model and fails to truly judge whether the semantics of the two are consistent. There may be a situation where the semantic similarity is low but the semantics are actually consistent; the other is the lack of consideration of the background text, and there may be a situation where although the semantics are inconsistent, they refer to the same thing in the text, resulting in misjudgment.
[0053] The relevant testing work lacks attention to the quality and coverage of generated test cases. Therefore, this application proposes a question-and-answer software testing method based on large language models. Using large language models as tools and aiming at answers, a method for generating a large number of high-quality test cases based on background text is proposed, which improves the testing method and theory of closed-world question-and-answer software. In the test case generation part, this application aims to provide a method for automatically generating a large number of high-quality test cases based on the original data set. By using large language model tools and designing various strict screening rules, the test cases are made more realistic, more capable of detecting errors, and more reliable, thereby improving the testing efficiency and solving the problem that the actual value of the existing work's test cases is not high. In the defect detection part, this application incorporates the internal semantics of the reference text into the comparison standard, aiming to provide a consistency comparison method that can not only accurately capture a large number of defects but also reduce false defect reports, making up for the defects of the existing work's defect detection methods that only consider isolated semantics.
[0054] As Figure 2 shown, this application mainly includes 5 modules. First, the information extraction module uses the technologies of large language models and natural language processing to extract relationship information and entity information from the background text of the seed data set as the standard answers for generating test cases. Second, the question generation module uses the information obtained in the first module as the answer and designs prompt words containing 4 aspects of content, allowing the large language model to generate questions based on the background text. Third, the result screening module designs two aspects of screening rules to screen the questions generated by the large language model in the previous module to obtain the test cases of this method. Fourth, execute the test cases on the software to be tested to obtain the answers output by the software to be tested. Finally, the defect detection module designs a two-aspect answer consistency comparison method to output the defects found by this application in this question-and-answer software.
[0055] Specifically, it includes the following steps:
[0056] 1. Information extraction
[0057] This step is to extract information from the background text as the answer for generating questions. Specifically, in order to efficiently utilize the information contained in the background text and better control question generation, this application pre-establishes two types of information: entity information (nouns, verbs, noun phrases, verb phrases) and relationship information (relationship triples) as the targets for information extraction. For the entity information part, taking sentences as units, the NLP toolkit stanza developed by Stanford University is used for processing. For gerund phrases and gerund words, first use stanza for part-of-speech and constituent analysis for preliminary extraction, and then set some screening conditions according to the dependency relationship of the sentences for screening, and output the extracted grammatical information. The screening rules here are:
[0058] Phrase:
[0059] 1) The length should not be less than 2
[0060] 2) If the phrase obtained from component analysis contains clause elements, the clauses therein shall be removed
[0061] 3) Duplicate removal
[0062] Words:
[0063] 1) A noun cannot be a pronoun, a verb cannot be a non-finite verb, and it cannot play a modifying or infinitive role in a dependency relationship
[0064] 2) A word cannot be included in the phrase extracted from this sentence
[0065] For the relationship information part, the extraction target is a relationship triple in the format of [entity 1, relationship, entity 2]. Since the effect of existing relationship extraction tools is poor, this application selects the latest large language model ChatGPT as the relationship extraction tool. This application combines the prompting engineering techniques of multi-instance and chain of thought to construct the prompt for relationship extraction. Based on this prompt, the initial relationship triple is obtained from ChatGPT. After screening the results with some heuristic screening rules, the final relationship information triple is obtained. The screening rules are as follows:
[0066] 1) Both entities should be included in the explanation and the text, and the entity should be a noun or a noun phrase
[0067] 2) The relationship cannot be a non-finite verb
[0068] 2. Question generation based on large language models
[0069] In this step, questions will be generated based on the answers as test cases. To make the generated questions more realistic and diverse so as to discover more defects, ChatGPT is used as the question generation tool. Specifically, the information obtained in the first step is used as the answer, and the sentence where it is located and the background text are sent into the large model, and the large model is made to ask questions based on the answer to generate questions. To generate questions more efficiently, this application selects to construct the prompt in the format of "instruction + restriction + multi-instance + chain of thought".
[0070] 3. Result screening
[0071] Since the output of the large language model is random and may generate questions that do not meet the requirements, this application designs a complete evaluation link to screen the results. Specifically, the screening is divided into two aspects. On the one hand, heuristic rules are designed to screen the results generated by the large model, and the rules are set as
[0072] 1) The answer should not be a pronoun or a non-finite verb
[0073] 2) The answer should not appear in the question.
[0074] 3) For entity information questions, the answer should appear in the original sentence along with the reason; for relationship information questions, all triple information should appear in the explanation.
[0075] On the other hand, for the generated questions and their answers, obviously if it is correct, then the answer when the large model re - answers this question based on its background text should be the same as the standard answer. Therefore, this application constructs a prompt for the large model to answer questions, and in order to eliminate the large model's bias, the method of voting by raising hands is used to select the most frequently occurring result from the 5 outputs of the large model as the true result, and SIMCSE is used to calculate its semantic similarity with the standard answer, and those with a similarity greater than the set threshold are retained.
[0076] 4. Test case execution
[0077] This step runs the test cases obtained in the previous step on the software under test to get the answers of the software under test.
[0078] 5. Defect detection
[0079] This step aims to detect whether the test cases have discovered defects in the Q&A software under test. The specific method is to compare whether the standard answer in the test case is the same as the answer given by the software under test. First, consider the semantic consistency of the answers, and use SIMCSE to calculate the semantic similarity between the two. When the semantic similarity does not reach the preset threshold, further consider the diverse expressions in the background text to reduce the possibility of false positives. Considering that the same thing may have different expressions in the background text, ChatGPT is used as a tool to construct specific prompts, and the large model is used to score the consistency of the two answers based on the background text. Utilize the excellent reasoning ability of the large language model to ensure the accuracy of the detected defects, and at the same time reduce the risk of false positives caused by language expression differences. If the result with a score lower than the threshold is obtained, it is considered that the answers are inconsistent, indicating that the test case has discovered a defect.
[0080] The Q&A software testing method based on large language models provided by the embodiments of this application comprehensively uses various techniques in prompt engineering to construct effective prompts, generates a large number of natural test cases based on the information in the background text of the original dataset, and fully screens the generated results to ensure the quality of the test cases. Compared with existing testing methods, the test cases generated by this method are more natural and can trigger defects more effectively, thus improving the testing efficiency. This application improves the semantic consistency comparison method in the defect detection link, incorporates the background text semantics into the comparison considerations, uses a large language model as an auxiliary tool, and designs a set of comparison methods from two aspects: semantics + text semantics, which improves the accuracy of the consistency comparison link in the testing method and solves the problem of high false positive rate in existing methods.
[0081] Corresponding to the Q&A software testing method based on large language models described in the above embodiments, as Figure 3 shown, the embodiments of this application also provide a Q&A software testing device based on large language models. The Q&A software testing device based on large language models includes:
[0082] An information extraction module, configured to extract information from the background text of the existing dataset as the standard answer for generating questions;
[0083] A test case generation module, configured to use the large language model to ask questions based on the standard answer and generate questions as test cases;
[0084] A software testing module, configured to run the test cases on the Q&A software to be tested to obtain the answers of the Q&A software to be tested.
[0085] It should be noted that for the information interaction, execution process, etc. between the above modules / units, since they are based on the same concept as the method embodiments of this application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0086] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0087] An embodiment of this application also provides an electronic device, such as Figure 4 shown, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the question-answering software testing method based on a large language model provided in the first aspect.
[0088] In applications, an electronic device may include, but is not limited to, a processor and a memory. Figure 4 This is only an example of an electronic device and does not limit the electronic device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, input / output devices, network access devices, etc. The input / output devices may include cameras, audio acquisition / playback devices, displays, etc. The network access device may include a network module for wireless communication with external devices.
[0089] In applications, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0090] In an application, in some embodiments, the memory may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory may also be an external storage device of the electronic device, for example, a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory may also include both an internal storage unit and an external storage device of the electronic device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of computer programs. The memory may also be used to temporarily store data that has been output or is to be output.
[0091] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor can implement the steps in the above-mentioned method embodiments.
[0092] To implement all or part of the processes in the above method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program codes, and the computer program codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program codes to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0093] Those of ordinary skill in the art can realize that the device and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0094] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. Additionally, the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces, and the devices are indirectly coupled or communication-connected, which can be in electrical, mechanical or other forms.
[0095] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for testing a question-and-answer software based on a large language model, characterized in that, including: extracting information from the background text of the existing dataset as the standard answer for generating questions; using a large language model to ask questions based on the standard answer to generate questions as test cases; running the test cases on the question-and-answer software to be tested to obtain the answers of the question-and-answer software to be tested.
2. The method for testing question and answer software based on a large language model according to claim 1, wherein The extracting information from the background text of the existing dataset includes: extracting entity information and relationship information from the background text of the existing dataset, where the entity information includes nouns, verbs, noun phrases, and verb phrases, and the relationship information includes relationship triples.
3. The method for testing question and answer software based on a large language model according to claim 2, wherein The extracting entity information and relationship information from the background text of the existing dataset includes: performing part-of-speech and constituent analysis on gerund phrases and gerunds, setting screening conditions according to the dependency relationship of sentences for screening, and outputting the extracted grammatical information; constructing prompt words for relationship extraction by combining multiple examples and chain of thought, and based on the prompt words, using a large language model to obtain initial relationship triples.
4. The method for testing a question-and-answer software based on a large language model according to claim 1, wherein, The using a large language model to ask questions based on the standard answer to generate questions as test cases includes: inputting the standard answer combined with the background text into the large language model, and constructing prompt words in the format of instruction + restriction + multiple examples + chain of thought to generate questions as test cases.
5. The method for testing question-and-answer software based on a large language model according to claim 1, wherein After the using a large language model to ask questions based on the standard answer to generate questions as test cases, it further includes: screening the questions generated by the large language model by designing rules and re-asking the large language model.
6. The method for testing a question-and-answer software based on a large language model according to claim 1, wherein After the running the test cases on the question-and-answer software to be tested to obtain the answers of the question-and-answer software to be tested, it further includes: comparing whether the standard answer is consistent with the answer of the question-and-answer software to be tested to detect whether the test cases have discovered defects in the question-and-answer software to be tested.
7. The method for testing a question-and-answer software based on a large language model according to claim 6, wherein, The comparing whether the standard answer is consistent with the answer of the question-and-answer software to be tested to detect whether the test cases have discovered defects in the question-and-answer software to be tested includes: calculating the semantic similarity between the standard answer and the answer of the question-and-answer software to be tested; if the semantic similarity does not reach the first preset threshold, using a large language model to perform a consistency score on the two answers based on the background text; if the score is lower than the second preset threshold, determining that the test cases have discovered defects in the question-and-answer software to be tested.
8. An apparatus for testing a question-and-answer software based on a large language model, characterized in that, including: an information extraction module for extracting information from the background text of the existing dataset as the standard answer for generating questions; a use case generation module for using a large language model to ask questions based on the standard answer to generate questions as test cases; a software testing module for running the test cases on the question-and-answer software to be tested to obtain the answers of the question-and-answer software to be tested.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the question-and-answer software testing method based on a large language model according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the question-and-answer software testing method based on a large language model according to any one of claims 1 to 7.
Citation Information
Cited By
Question and answer method and device based on knowledge graph retrieval and large language model
CN120950652A