Question and answer agent testing method and device and electronic equipment

Through the Q&A agent testing method, the big model is used to compare the consistency of the output answers and the standard answers, and automatically generate and evaluate the test problem set, solving the existing agent test low degree of automation, and achieving an efficient and safe agent testing process.

CN120371698APending Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450841.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing agent test links have low degree of automation and low efficiency in relying on manual code writing, resulting in long test cycles and high cost, which affects the speed of R&D iteration.

Method used

The test method of Q&A agent is adopted. By inputting the test question set into the Q&A agent, the output answer set is obtained, and the standard answer set is obtained from the external database platform. The large model is used to compare consistency to obtain the test results. The test agent automatically generates the test question set and performs automated evaluation.

Benefits of technology

It realizes automation and efficiency of agent testing, lowers the test threshold, shortens the test cycle, improves the test efficiency, and enhances data security and the accuracy of test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371698A_ABST
    Figure CN120371698A_ABST
Patent Text Reader

Abstract

The invention provides a question and answer agent testing method, and relates to the technical field of testing, in particular to the technical field of artificial intelligence. According to the specific implementation scheme, a test question set is input into a to-be-tested question and answer agent, and an output answer set output by the question and answer agent based on the test question set is obtained; obtaining a standard answer set corresponding to the test question set from an external database platform; and comparing the consistency of the output answer set and the standard answer set through a large model, and obtaining a test result of the question and answer agent according to a consistency comparison result. According to the invention, automation of intelligent agent testing can be realized, and the testing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of testing technologies, and in particular to the field of artificial intelligence technologies. Background Art

[0002] With the rapid development of artificial intelligence technologies, intelligent agents have been widely applied in many fields such as smart home, intelligent customer service, and intelligent driving. However, there is a problem of poor automation in the existing intelligent agent testing process.

[0003] Currently, intelligent agent testing mainly relies on manually writing code. On the one hand, testers need to have professional programming knowledge and design complex test codes for different intelligent agents, which requires extremely high professional qualities of testers. On the other hand, manually writing code is inefficient, and the entire testing process takes a long time, resulting in a significant extension of the testing cycle, an increase in R & D costs, and seriously affecting the R & D iteration speed of intelligent agents. In this context, there is an urgent need for an efficient and convenient intelligent agent testing method to achieve the intelligence and high efficiency of intelligent agent testing. Summary of the Invention

[0004] The present disclosure provides a method, device, and electronic device for mining risk users to solve at least one of the above technical problems.

[0005] According to one aspect of the present disclosure, there is provided a method for testing a question - answering intelligent agent, wherein the method is applied to a testing intelligent agent, and the method includes:

[0006] Inputting a test question set into the question - answering intelligent agent to be tested, and obtaining an output answer set output by the question - answering intelligent agent based on the test question set;

[0007] Obtaining a standard answer set corresponding to the test question set from an external database platform;

[0008] Comparing the consistency between the output answer set and the standard answer set through a large model, and obtaining a test result of the question - answering intelligent agent according to the consistency comparison result.

[0009] According to another aspect of the present disclosure, there is provided a device for testing a question - answering intelligent agent, wherein the device includes:

[0010] An output answer module, configured to input a test question set into the question - answering intelligent agent to be tested, and obtain an output answer set output by the question - answering intelligent agent based on the test question set;

[0011] A standard answer module, configured to obtain a standard answer set corresponding to the test question set from an external database platform;

[0012] A comparison test module for comparing the consistency between the output answer set and the standard answer set through a large model, and obtaining the test result of the Q&A agent according to the consistency comparison result.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.

[0018] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the above method when executed by a processor.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0021] Figure 1 is a schematic structural diagram of the application scenario framework of the present disclosure;

[0022] Figure 2 is a schematic flowchart of a Q&A agent test method provided by the first embodiment of the present disclosure;

[0023] Figure 3 is a schematic flowchart of an exemplary process of generating a test question set;

[0024] Figure 4 is a schematic flowchart of an exemplary S102;

[0025] Figure 5 is a schematic flowchart of an exemplary S1022;

[0026] Figure 6 is a schematic flowchart of an exemplary S103;

[0027] Figure 7 It is a schematic flowchart of a question-and-answer agent testing device according to the second embodiment of the present disclosure;

[0028] Figure 8 It is a block diagram of an electronic device for implementing the method according to the embodiment of the present disclosure. Detailed implementation manners

[0029] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0030] In the case of no conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0031] As used herein, the term "and / or" includes any and all combinations of one or more related listed items.

[0032] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0033] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless clearly defined herein.

[0034] The question-and-answer agent testing method according to the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, a display device (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in a memory. Alternatively, the question-and-answer agent testing method provided by the present disclosure can be executed by a server.

[0035] See Figure 1 , Figure 1Schematic diagram of the structure of the application scenario framework of the present disclosure. The figure includes a test agent 1, a question-and-answer agent 2, and an external database platform 3. The present disclosure is applied to the test agent 1, which is used to test the question-and-answer agent 2. The test agent 1 establishes communication with the external database platform 3 to query the standard answers to the test questions through the external database platform 3. Among them, there can be multiple external database platforms, corresponding to data queries for different test questions, which are not limited here.

[0036] In the first embodiment of the disclosure, participate Figure 2 , Figure 2 The flowchart shows a method for testing a question-and-answer agent provided by the first embodiment of the present disclosure. This method is applied to a test agent, and the method includes:

[0037] S101. Input a test question set into the question-and-answer agent to be tested, and obtain an output answer set output by the question-and-answer agent based on the test question set.

[0038] S102. Obtain a standard answer set corresponding to the test question set from the external database platform.

[0039] S103. Compare the consistency between the output answer set and the standard answer set through a large model, and obtain the test result of the question-and-answer agent according to the consistency comparison result.

[0040] Agent: Refers to a system that has the ability to autonomously sense the environment, process information, and make decisions and actions based on preset rules or learning experiences. It can automatically complete complex tasks without real-time human intervention.

[0041] In the scenario of the present disclosure, the test agent, as a specific type of agent, is given the ability to automatically test the question-and-answer agent. In some examples, before the test starts, the test agent can automatically generate a test question set and guide the question-and-answer agent to be tested to generate an output answer set. Of course, in some examples, the test question set can also be generated externally and then imported into the test agent, which is not limited here. Moreover, the test agent can also obtain the standard answer set from the external database platform, and use a large model to compare and analyze the output answer set and the standard answer set to complete the test evaluation of the question-and-answer agent. Among them, the question-and-answer agent to be tested can be used to respond to various questions and is an intelligent system that simulates human dialogue logic and provides accurate answers.

[0042] As an example, in order to cover the question-and-answer tests in multiple knowledge fields, the test question set generated by the test agent can cover various question types. For example, the test question set generated by the test agent is a test question set in the transportation field, a test question set in the medical field, etc., which are not limited here.

[0043] The method provided by the present disclosure inputs a test question set into the question-and-answer intelligent agent to be tested, obtains an output answer set output by the question-and-answer intelligent agent based on the test question set, obtains a standard answer set corresponding to the test question set from an external database platform, compares the consistency between the output answer set and the standard answer set through a large model, and obtains the test result of the question-and-answer intelligent agent according to the consistency comparison result. Thus, the test of the question-and-answer intelligent agent is completed by using the test intelligent agent, without the need for testers to write test codes and conduct analysis, which solves the problem of manually writing complex test codes and reduces the test threshold. Moreover, it can greatly improve the test efficiency. By automating the test process, steps such as generating test questions, obtaining standard answers, and comparing answer consistency are integrated, greatly shortening the test cycle and accelerating the R & D iteration speed of the question-and-answer intelligent agent. Further, the method provided by the present disclosure encapsulates the test process in the test intelligent agent, reducing the risk of data leakage during the test and exposing the prompt words used for testing. Compared with other test methods, it can effectively protect data security, ensure the stability and reliability of the test process, and further make the test of the question-and-answer intelligent agent more intelligent, secure and efficient.

[0044] In some examples, refer to Figure 3 , Figure 3 shows a schematic flowchart of an exemplary process for generating a test question set. In the related technical field, it takes a long time for manual labor to produce a large number of test questions. In the present disclosure, the test intelligent agent can automatically generate a test question set, thus greatly improving the test efficiency. Specifically, the steps for generating the test question set are as follows:

[0045] S001. Automatically input a preset test question prompt and item information in multiple dimensions into the large model to obtain a test question set output by the large model.

[0046] Among them, the test question prompt (prompt) is used to instruct the large model to randomly permute and combine at least two items in the item information in multiple dimensions according to the preset priority of each dimension to generate a test question set.

[0047] When the test process starts, the test intelligent agent will automatically generate a test question set.

[0048] Among them, the item information is the basic information for generating test questions. For example, taking the test question set as a test question set in the transportation field, the item information in multiple dimensions includes at least two of the following: intersection item information, traffic event type item information, query time item information, query index item information, etc.; taking the test question set as a test question set in the medical field, the item information in multiple dimensions includes at least two of the following: disease type item information, treatment method item information, treatment cycle item information, etc. Of course, it can also include more types of item information, which is not limited here.

[0049] Exemplarily, any item information may include multiple sub-information. The intersection item information may be in the form of an intersection item list, including multiple sub-information of intersection names; the traffic event type item information may be a traffic time type list, and the traffic time type list includes multiple sub-information of traffic event types, such as traffic accident event type sub-information, road construction event type sub-information, bad weather event type sub-information, motor vehicle reverse event type sub-information, signal light failure event type sub-information, etc., which are not limited here; the query time item information may include multiple sub-information, and the multiple sub-information are, for example, at least one of hour level, day level, week level, month level, etc., which are not limited here; the query indicator item information may include multiple sub-information, such as at least one of quantity type sub-information, percentage type sub-information, etc., which are not limited here. Through the item information of multiple dimensions, basic information for comprehensively constructing a test question set is provided for the large model. The test question prompt words can perform permutation and combination on at least two of the item information of multiple dimensions to generate a test question set. In some examples, in order to comprehensively cover various Jiangtong problems, permutation and combination can be performed on each item of the item information of multiple dimensions to generate a test question set, which is not limited here.

[0050] In S001, the test question prompt words are used to indicate the direction for the large model to output test questions and provide the task instructions for outputting test questions. For example, in the scenario of intelligent traffic condition query, the test question prompt words designed around the intersection item information, that is, the preset intersection item information has the highest priority, can guide the large model to generate test questions focusing on the intersection traffic conditions, making the questions closely fit the actual application scenario. In the traffic emergency command scenario, the test question prompt words centered on the traffic event type, that is, the preset traffic event type item information has the highest priority, will let the large model generate questions around events such as traffic accidents and road control, ensuring the accuracy of the test direction. All in all, the information carried by the test question prompt words can activate the logical reasoning mechanism of the large model to guide the large model to perform correlation analysis on each item in the item information of each dimension according to the preset priorities of each dimension, permute and combine each item of different dimensions to generate a series of reasonable and high-quality test questions, and combine them into a test question set.

[0051] On the one hand, by reasonably designing the test question prompts, the large model can be made to generate test questions covering various scenarios. Since the test question prompts can guide the large model to combine multi-dimensional information, prompting the test questions to comprehensively cover different time periods, different types of traffic events, different intersection locations, and different query metrics, avoiding omissions in testing and enhancing the comprehensiveness of testing. On the other hand, by generating the test question set in this way, it is possible to avoid the large model groping blindly during the test question generation process, reduce unnecessary consumption of computing resources, save time costs, greatly improve the generation efficiency of test questions, and thereby improve the efficiency of the entire test process.

[0052] Furthermore, after receiving the test question prompt and project information in multiple dimensions, the large model, based on the powerful semantic understanding ability of the Transformer architecture, uses the multi-head attention mechanism to perform refined feature extraction on the input information, and then combines and optimizes the features. The test question prompt guides the large model to perform deep random permutations and combinations on each item of the traffic project information in multiple dimensions according to the preset priorities of the project information in each dimension. Taking the intelligent traffic flow monitoring scenario as an example, assume that the project information in multiple dimensions includes intersection project information, traffic event type project information, query time project information, and query indicator project information, and assume that the priorities are sorted from high to low as intersection project information - traffic event type project information - query time project information - query indicator project information. Then the large model formulates a basic question framework of "a certain intersection - a certain traffic time type - a certain period - how many times", preferentially selects any item in the intersections included in the intersection project information, and successively combines it with any item in the traffic event type project information, any item in the query time project information, and any item in the query indicator project information to generate multiple test questions. Assume that the priorities are sorted from high to low as traffic event type project information - query time project information - intersection project information - query indicator project information. Then the large model formulates a basic question framework of "a certain traffic time type - a certain period - a certain intersection - how many times", preferentially selects any item in the traffic time types included in the traffic event type project information, and successively combines it with any item in the query time project information, any item in the intersection project information, and any item in the query indicator project information to generate multiple test questions. The above is only an example. Of course, the test question set is not limited to the application scenarios in the traffic field and can also be applied to the application scenarios in other fields, which are not limited here. Through this multi-dimensional and deep random permutation and combination, not only can the homogenization of test questions be effectively avoided, making the generated test question set richer and more diverse in structure and content, but also it can ensure that the test question set comprehensively covers various situations that the question-and-answer agent may encounter in the actual application scenario. Through this question generation mechanism, the accuracy and effectiveness of the test process can be greatly improved, providing a basis for accurately evaluating the performance of the question-and-answer agent in the future, ensuring the accuracy of the test for the question-and-answer agent, and in this way, the steps of manually configuring the test question set can be omitted, and the test question set can be automatically generated, effectively improving the generation efficiency of the test question set.

[0053] In some examples, the project information in multiple dimensions in S001 is obtained in the following way:

[0054] Through the Retrieval-Augmented Generation (RAG) mechanism, retrieve and obtain the project information in multiple dimensions from an external knowledge base.

[0055] The RAG mechanism is an artificial intelligence technology that combines information retrieval technology with a text generation model. In this disclosure, a large model can be used as the text generation model. Traditional text generation models rely on the knowledge learned during pre-training and it is difficult to obtain the latest and most accurate information. The RAG mechanism allows the model to retrieve relevant information from an external knowledge base when generating content, so as to guide the generation of text. In this disclosure, the RAG mechanism connects the test agent with the external knowledge base, providing a more diverse and accurate data source for the test process. Inside the test agent, a stable connection can be established with the external knowledge base through the Hypertext Transfer Protocol (HTTP) or other standardized interface protocols. When the test process starts, before the test agent inputs data into the large model, first, through the RAG mechanism, based on the understanding of the test scenario and objectives, a retrieval request for project information in a certain dimension is generated. For example, in the intelligent transportation scenario, a retrieval request containing the intersection name can be generated, such as "Retrieve the [intersection name] available in [a certain area]", so as to generate intersection project information according to the queried intersection names; a retrieval instruction containing traffic events can be generated, such as "Retrieve the traffic events and their types that occurred in [a certain area]", so as to generate traffic event type project information containing classified traffic events according to the queried traffic events and types. In some examples, the project information can be generated as a project list.

[0056] For example, the above retrieval request is sent to the external knowledge base, and the external knowledge base parses the request according to the preset indexing rules and data storage structure. Through efficient retrieval algorithms such as inverted index and vector retrieval, the matching information can be quickly located in the massive knowledge base data (such as traffic data). Taking traffic data as an example, the knowledge base stores the data in layers according to dimensions such as intersections, traffic event types, and time, and can quickly respond to retrieval requests regarding specific intersections, specific types of traffic events, and traffic flow during specific time periods. When the matching information is retrieved, the knowledge base packages the data and feeds it back to the test agent through the established interface. The test agent cleans and preprocesses the returned data, removing redundant and incorrect data to ensure the accuracy and availability of the data. Subsequently, project information in different dimensions is generated according to the processed data and transmitted to the input end of the large model together with the preset test question prompt words.

[0057] Obtaining multi-dimensional project information from the external knowledge base through the RAG mechanism can ensure the timeliness and comprehensiveness of the project information. Moreover, this method can effectively reduce the data storage pressure of the test agent and optimize the system performance.

[0058] In some examples, see Figure 4 , Figure 4Shows a schematic flowchart of an exemplary S102. S102 includes:

[0059] S1021. Input each test question included in the test question set into the large model to obtain the questioning parameters output by the large model.

[0060] In some examples, the questioning parameters are in the form of machine language. In other words, in this case, S1021 is: Input each test question included in the test question set into the large model to obtain the questioning parameters output by the large model and in the form of machine language. In this way, it is convenient to directly use the questioning parameters in the form of machine language to generate query statements and query the standard answers corresponding to the test questions from an external platform.

[0061] In some examples, before S101, the method provided by the present disclosure further includes:

[0062] Input each test question sample and its corresponding preset parameter prompt words into the large model to obtain the questioning parameters output by the large model and corresponding to each test question sample.

[0063] Among them, the parameter prompt words are used to indicate the parameter fields included in the questioning parameters generated by the large model, and can also be used to indicate at least one of the format of the parameter fields included in the questioning parameters and the value range of the parameter fields.

[0064] In this way, to enable the large model to learn the ability to convert test questions into questioning parameters, during the training phase, a rich training data set can be constructed. The training data set includes a large number of pairs of test question samples and corresponding parameter prompt words, as well as the expected output questioning parameters. Based on the powerful self-attention mechanism and multi-layer neural network of the Transformer architecture, the large model realizes the understanding and analysis of test questions and parameter prompt words, so as to learn the ability to convert various test questions into questioning parameters.

[0065] During the training process, the large model receives the test question sample and the parameter prompt words as inputs, performs word embedding and position encoding on the input text through the Transformer architecture, and converts the questions of the test question sample into a vector representation that can be processed by the model. With the help of the self-attention mechanism, the multi-layer neural network of the model captures the semantic associations between different words in the test question sample and the parameter prompt words, and deeply understands the question intention and parameter requirements. Based on these understandings, the model generates questioning parameters through operations. Compare the generated questioning parameters with the expected output questioning parameters in the training data set, and calculate the loss value between the two. With the help of the backpropagation algorithm, the model continuously adjusts the parameters in the neural network to reduce the loss value and optimize the output result of the model. As the training continues, the large model gradually learns to accurately generate the questioning parameters that meet the requirements according to different test question samples and parameter prompt words.

[0066] As an example, in a traffic scenario, a training sample might be: The test question is "Get the [number] of [motor vehicle red light running incidents] that occurred at [Intersection A] on [January 2, 2024]". The parameter prompt words stipulate that the query parameters should include fields such as the start time field, the end time field, the traffic event name field, the intersection name field, etc. Exemplarily, it can further specify that the format of the intersection name field is to use the full name, the time formats of the start time field and the end time field, etc. Further, it can also indicate the file format of the query parameters expected to be output, for example: a text-based data interchange format (JavaScript Object Notation, JSON) file. Thus, based on the example of the above training sample, the query parameters finally output by the large model can be encapsulated in the form of a JSON file as follows:

[0067] {

[0068] "startTime": "2024-01-02 00:00:00"

[0069] "endTime": "2024-01-02 23:59:59"

[0070] "eventName": ["motor vehicle red light running incidents"]

[0071] "roadName": "Intersection A"

[0072] }

[0073] Among them, "startTime" is the start time field; "endTime" is the end time field; "evenName" is the traffic event name field; "roadName" is the intersection name field.

[0074] The above is only an example of converting a test question sample into query parameters and does not limit this application.

[0075] In some examples, the query parameters are encapsulated in a JSON file. Using the JSON format to define the query parameters has good readability and generality, enabling different systems and programming languages to easily parse and generate them. In particular, query statements can be directly generated based on the parameter fields included in the query parameters of the JSON file, facilitating querying from an external database, and omitting the step of manually converting the test question in natural semantics into query parameters in machine language form, thus ensuring the efficiency of data interaction between the test agent and the external database platform.

[0076] S1022. Obtain the standard answer set corresponding to each question parameter from the corresponding external database platform based on each question parameter.

[0077] After the test agent obtains the question parameters output by the large model, it will convert them into query statements that can be recognized by the external database platform. The external database platform stores a large amount of accurate data covering multi-dimensional information. The external database efficiently stores and manages data according to its own data architecture and index rules. When the external database platform receives the query statements sent by the test agent, it will quickly perform retrieval and matching based on data indexing and query algorithms, and accurately locate the standard answers corresponding to the question parameters. After retrieving the matching information, the external database platform will organize these standard answers into a standard answer set and feedback it to the test agent.

[0078] In some examples, refer to Figure 5 , Figure 5 shows a schematic flow diagram of an exemplary S1022. S1022 includes:

[0079] S10221. Based on any question parameter, determine the corresponding target external data platform and its target interface.

[0080] After the test agent obtains the question parameters output by the large model, based on any question parameter, determine the corresponding target external data platform and its target interface to obtain the query destination.

[0081] In some examples, S10221 includes:

[0082] Sub-step one: Determine the target parameter fields included in the question parameter.

[0083] After the test agent converts the test question into question parameters in the form of machine language, it can parse the question parameters to determine the target parameter fields included in the question parameters. Taking the question parameters {"evenName": ["motor vehicle running a red light event"]; "roadName": "Intersection A"} as an example, the parameter parsing module identifies the target parameter fields included therein as: intersection name field and traffic event name field.

[0084] Sub-step two: In the preset interface lookup table, based on the target parameter fields, determine the corresponding target external data platform and the target interface corresponding to the target external data platform.

[0085] The test agent is preset with an interface lookup table, which is preset with the mapping relationships of multiple parameter fields - external data platforms - interfaces. After obtaining the target parameter field, the test agent performs matching in the interface lookup table. For example, it is stipulated in the lookup table that the target parameters include the intersection name field and the traffic event name field, and it should be docked with the traffic event statistics database platform, and its corresponding target interface is " / traffic_event_query". Based on this rule, the test agent determines that the target external data platform corresponding to the above-mentioned question parameters is the traffic event statistics database platform, and the target interface is " / traffic_event_query".

[0086] S10222. Generate a query statement that conforms to the target external data platform and send it to the target interface of the target external data platform to obtain the standard answer returned by the target external data platform in response to the query statement.

[0087] After determining the target external data platform and its target interface, the test agent needs to generate a query statement that meets the requirements of the target external data platform and send it to the target interface to obtain the standard answer. The test agent generates the corresponding query statement based on the interface specification of the target external data platform and the content of the question parameters. The query statement can adopt various formats, which are not limited here. Exemplarily, taking the traffic flow statistics database platform as an example, its interface specification requires that the query statement adopt the Structured Query Language (SQL) format. For the question parameters {"evenName": ["motor vehicle running a red light event"]; "roadName": "Intersection A"}, the following query statement can be generated: SELECT * FROM traffic_event_table; WHERE eventName ='motor vehicle running a red light event' AND roadName = 'Intersection A'.

[0088] The test agent sends the generated query statement to the " / traffic_event_query" interface of the traffic event statistics database platform and obtains the standard answer of the number of motor vehicle running a red light events occurring at Intersection A returned by the traffic event statistics database platform.

[0089] The above is only an example of generating a query statement and does not limit this application.

[0090] S10223. Combine the standard answers corresponding to the question parameters of each test question into a standard answer set.

[0091] It should be noted that in the disclosure, the step of "generating a query statement that conforms to the target external data platform" in S10222 can be executed by a large model or by a built-in plugin.

[0092] In an embodiment adopting a large model, during the training process, through learning a large number of pairs of question parameters and corresponding query instruction sample data, the large model masters the conversion rule from question parameters to query instructions, so as to accurately output query instructions in actual applications.

[0093] In an embodiment adopting a plug-in, by building in a plug-in for data query, the question parameters can be quickly converted into query statements that meet the requirements of the target external data platform. For example, adopting the CodeFuse plug-in, the CodeFuse plug-in supports multiple mainstream programming languages and can adapt to the query instruction syntax requirements of different external data platforms. Another example is the QueryMaster plug-in, which internally integrates a rich database interface specification and a query statement generation rule library, covering common relational databases and various professional field databases. Of course, other plug-ins can also be adopted, which are not limited here.

[0094] Generating query statements through two approaches of a large model and built-in plug-ins endows the question-answering intelligent body test system in the transportation field with stronger adaptability and scalability.

[0095] In some examples, refer to Figure 6 , Figure 6 which shows a schematic flowchart of an exemplary S103. S103 includes:

[0096] S1031. Input the output answer set and the standard answer set into the large model. The large model sequentially compares the consistency between each output answer in the output answer set and the corresponding standard answer in the standard answer set, and gives a single-item test question score.

[0097] The test intelligent body inputs the output answer set and the standard answer set into the large model for comparison. During the comparison process, the large model will sequentially process each output answer in the output answer set and the corresponding standard answer in the standard answer set, and conduct in-depth semantic analysis on each output answer and the corresponding standard answer. When judging the consistency of data, the large model will accurately compare specific data such as numerical values, names, and times in the text.

[0098] Exemplarily, when evaluating semantic consistency, the large model can understand the meaning expressed by the text and judge whether the two are equivalent at the semantic level. Taking the description of a traffic event as an example, the output answer states that "a rear-end collision occurred on a certain section of the road", and the standard answer is "a vehicle rear-end situation occurred on a certain section of the road". Although the expressions are slightly different, the large model can determine that the two are semantically consistent through semantic understanding. When evaluating data consistency, the large model can compare the difference between the data given by the two to determine data consistency.

[0099] In some examples, the large model also analyzes the semantic relevance between each test question and the corresponding output answer in the output answer set to determine whether the semantic understanding of the test question by the large model and the output answer given are semantically consistent.

[0100] Based on the preset scoring criteria, the large model quantitatively scores the consistency comparison results to give the score for a single test question.

[0101] In some examples, after S1031, the method provided by the present disclosure further includes:

[0102] The large model generates a result display text according to the consistency comparison results. Among them, the result display text may include: the output answer, the standard answer, the score for a single test question, and may also include at least one of the question-and-answer time of the question-and-answer agent and the reason for scoring, which is not limited here.

[0103] S1032. Synthesize the scores of multiple single test questions to obtain the test result of the question-and-answer agent.

[0104] After obtaining the scores of multiple single test questions, the test agent comprehensively analyzes and integrates the scores of multiple single test questions to obtain the final test result. Specifically, multiple algorithms can be used to obtain the test result of the question-and-answer agent, such as the weighted average method. According to the importance of different test questions, corresponding weights are assigned to the scores of each single test question, and then the weighted average value is calculated to obtain the final test result.

[0105] By comparing this test result with the pre-set performance standard, it can be clearly judged whether the question-and-answer agent meets the expected requirements. If the test result is higher than the standard score line, it indicates that the performance of the question-and-answer agent meets the standard; otherwise, the question-and-answer agent needs to be optimized and improved.

[0106] In some examples, the consistency between the output answer set and the standard answer set includes data consistency and / or semantic consistency.

[0107] In some examples, the consistency between the output answer set and the standard answer set covers data consistency and / or semantic consistency. This multi-dimensional consistency evaluation enables the test system to more comprehensively and accurately measure the performance of the question-and-answer agent, avoiding one-sidedness caused by single-dimensional evaluation.

[0108] Through S1031 and S1032, an efficient evaluation mechanism is constructed, and through the analysis ability of the large model, the performance of the question-and-answer agent can be accurately evaluated.

[0109] In some examples, the method provided by the present disclosure further includes:

[0110] Record the test process data;

[0111] The test process data includes at least one of the following: the generated test question set, the output answer set output by the Q&A agent, the standard answer set obtained from an external database platform, the consistency comparison result between the output answer set and the standard answer set, and the test result.

[0112] Among them, the steps of recording the test process data include storing the test process data in a specified storage location of a local storage device or a remote server for subsequent query and analysis, realizing the persistence of test data, and ensuring data security.

[0113] In summary, the method provided by the present disclosure automatically completes the entire process from test question generation, answer set comparison to test result output. Users do not need to configure the code by themselves, reducing the time consumption caused by operations such as code deployment and environment configuration, significantly shortening the test cycle, and improving the test efficiency. Moreover, using the test agent to execute the test without exposing the code and prompt words greatly reduces the security risks caused by code and prompt word leakage. It avoids malicious users from attacking or misusing the test system using open-source code and prompt words, effectively ensuring the security of test data and test processes. The test agent can autonomously coordinate the work of each module according to preset rules and goals, realizing the automation and intelligence of the test process.

[0114] In the second embodiment of the disclosure, refer to Figure 7 , Figure 7 which shows a schematic flowchart of a Q&A agent test device according to the second embodiment of the present disclosure. The device includes:

[0115] An output answer module 701, configured to input the test question set into the Q&A agent to be tested, and obtain an output answer set output by the Q&A agent based on the test question set;

[0116] A standard answer module 702, configured to obtain a standard answer set corresponding to the test question set from an external database platform;

[0117] A comparison test module 703, configured to compare the consistency between the output answer set and the standard answer set through a large model, and obtain the test result of the Q&A agent according to the consistency comparison result.

[0118] In some examples, the standard answer module 702 is specifically configured to:

[0119] Input each test question included in the test question set into the large model to obtain the question parameters output by the large model;

[0120] Based on each question parameter, obtain a standard answer set corresponding to each question parameter from the corresponding external database platform.

[0121] In some examples, the standard answer module 702 is specifically used when obtaining the standard answer sets corresponding to each question parameter from the corresponding external database platform based on each question parameter as follows:

[0122] Based on any question parameter, determine the corresponding target external data platform and its target interface;

[0123] Generate a query statement that conforms to the target external data platform and send it to the target interface of the target external data platform to obtain the standard answer returned by the target external data platform in response to the query statement;

[0124] Combine the standard answers corresponding to the question parameters of each test question into a standard answer set.

[0125] In some examples, when the standard answer module 702 determines the corresponding target external data platform and its target interface based on any question parameter, it is specifically used for:

[0126] Determine the target parameter fields included in the question parameter;

[0127] In the preset interface lookup table, determine the corresponding target external data platform based on the target parameter field, and the target interface corresponding to the target external data platform.

[0128] In some examples, the question parameter is encapsulated as a data exchange format file based on text.

[0129] In some examples, the steps to generate the test question set are as follows:

[0130] Automatically input the preset test question prompt words and project information in multiple dimensions into the large model to obtain the test question set output by the large model;

[0131] The test question prompt words are used to instruct the large model to randomly permute and combine at least two of the project information in multiple dimensions according to the priorities of each preset dimension to generate the test question set.

[0132] In some examples, the test question set is a test question set in the traffic field;

[0133] The project information in multiple dimensions includes at least two of the following: intersection project information, traffic event type project information, query time project information, query index project information.

[0134] In some examples, the project information in multiple dimensions is obtained through the following method:

[0135] Through the retrieval enhancement generation mechanism, retrieve and obtain the project information in multiple dimensions from the external knowledge base.

[0136] In some examples, the comparison test module 703 is specifically used for:

[0137] The output answer set and the standard answer set are input into the large model, and the large model sequentially compares the consistency between each output answer in the output answer set and the corresponding standard answer in the standard answer set, and gives a score for the single test question.

[0138] By synthesizing the scores of multiple single test questions, the test result of the Q&A agent is obtained.

[0139] In some examples, the consistency between the output answer set and the standard answer set includes data consistency and / or semantic consistency.

[0140] In some examples, the device further includes: a recording module for recording test process data;

[0141] The test process data includes at least one of the following:

[0142] The generated test question set;

[0143] The output answer set output by the Q&A agent;

[0144] The standard answer set obtained from an external database platform;

[0145] The comparison result of the consistency between the output answer set and the standard answer set;

[0146] The test result.

[0147] According to an embodiment of the present disclosure, the present disclosure further provides a test agent for artificial intelligence, which is configured to execute the various methods described above, such as the Q&A agent test method.

[0148] Exemplarily, the test agent may include an input module, a processing module, and an output module.

[0149] The input module is used to receive input information;

[0150] The processing module is used to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the Q&A agent test method provided according to the embodiment of the present disclosure by calling the large model.

[0151] The output module is used to output the output information obtained by the processing module.

[0152] According to an embodiment of the present disclosure, the input module is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (such as users or the external environment), and converting it into a format that the test agent can understand and process. The input module is the primary link for the test agent to interact with the outside world, enabling the test agent to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0153] In the example, the input module can input the output answer set, standard answer set, etc. described above.

[0154] In the example, the processing module is the core support for testing the ability of the intelligent agent to handle complex tasks. The processing module can execute the question-and-answer intelligent agent testing method described above.

[0155] In the example, the performance of the processing module can be closely related to the large model on which the intelligent agent is based. In order to fully utilize the capabilities of the large model, the internal structure of the processing module can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.

[0156] In the example, after the intelligent agent obtains the output answer set and the standard answer set, the processing module can utilize the consistency between the output answer set and the standard answer set of the large model, obtain the test result of the question-and-answer intelligent agent based on the consistency comparison result, and transmit the test result to the output module.

[0157] It can be understood that although large language models have excellent language understanding and generation capabilities, like humans, the tasks they can solve without any tools are very limited. When the intelligent agent is given the ability to call tools, it can achieve tasks such as performing mathematical operations with the help of a calculator, conducting data analysis with the help of Python, and obtaining weather forecasts with the help of a search engine.

[0158] In the example, the output module can output the test result of the question-and-answer intelligent agent described above.

[0159] The intelligent agent according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.

[0160] In the technical solution of the present disclosure, the processing of the user's personal information, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.

[0161] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the user's authorization or consent has been obtained.

[0162] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] Such as Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 802 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0164] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0165] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the question-and-answer agent test method. For example, in some embodiments, the question-and-answer agent test method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the question-and-answer agent test method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured with the question-and-answer agent test method in any other appropriate way (e.g., by means of firmware).

[0166] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0167] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0168] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0170] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0171] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0172] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0173] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for testing a question-and-answer intelligent agent, wherein, Including: Input the test question set into the Q&A intelligent agent to be tested, and obtain the output answer set output by the Q&A intelligent agent based on the test question set; Obtain the standard answer set corresponding to the test question set from an external database platform; Compare the consistency between the output answer set and the standard answer set through a large model, and obtain the test result of the Q&A intelligent agent according to the consistency comparison result.

2. The method according to claim 1, wherein The obtaining the standard answer set corresponding to the test question set from an external database platform includes: Input each test question included in the test question set into a large model, and obtain the question parameters output by the large model; Based on each of the question parameters, obtain the standard answer set corresponding to each of the question parameters from the corresponding external database platform.

3. The method according to claim 2, wherein, The obtaining the standard answer set corresponding to each of the question parameters from the corresponding external database platform includes: Based on any one of the question parameters, determine the corresponding target external data platform and its target interface; Generate a query statement that conforms to the target external data platform, and send it to the target interface of the target external data platform to obtain the standard answer returned by the target external data platform in response to the query statement; Combine the standard answers corresponding to the question parameters of each of the test questions into the standard answer set.

4. The method according to claim 3, wherein, The determining the corresponding target external data platform and its target interface based on any one of the question parameters includes: Determine the target parameter fields included in the question parameters; In a preset interface lookup table, determine the corresponding target external data platform and the target interface corresponding to the target external data platform based on the target parameter fields.

5. The method according to any one of claims 2-4, wherein, The question parameters are encapsulated into a data exchange format file based on text.

6. The method according to any one of claims 1-5, wherein, The steps of generating the test question set are as follows: Automatically input a preset test question prompt and project information in multiple dimensions into a large model, and obtain the test question set output by the large model; The test question prompt is used to instruct the large model to randomly permute and combine at least two of the project information in multiple dimensions according to the priorities of each dimension preset, so as to generate the test question set.

7. The method according to claim 6, wherein, The test question set is a test question set in the transportation field; The project information in multiple dimensions includes at least two of the following: intersection project information, traffic event type project information, query time project information, query index project information.

8. The method according to claim 6 or 7, wherein, The project information in multiple dimensions is obtained through the following method: Retrieve and obtain the project information in multiple dimensions from an external knowledge base through a retrieval enhancement generation mechanism.

9. The method according to any one of claims 1-8, wherein, The comparing the consistency between the output answer set and the standard answer set through a large model, and obtaining the test result of the Q&A intelligent agent according to the consistency comparison result includes: Input the output answer set and the standard answer set into a large model, and the large model sequentially compares the consistency between each output answer in the output answer set and each corresponding standard answer in the standard answer set, and gives a single test question score; Integrate multiple single test question scores to obtain the test result of the Q&A intelligent agent.

10. The method according to any one of claims 1-9, wherein, The consistency between the output answer set and the standard answer set includes data consistency and / or semantic consistency.

11. According to the method described in any one of claims 1-10, wherein, The method further includes: recording test process data; The test process data includes at least one of the following: The generated test question set; The output answer set output by the question-and-answer intelligent agent; The standard answer set obtained from the external database platform; The consistency comparison result between the output answer set and the standard answer set; The test result.

12. An intelligent agent test device for question and answer, wherein, The device includes: An output answer module, configured to input a test question set into a question-and-answer intelligent agent to be tested, and obtain an output answer set output by the question-and-answer intelligent agent based on the test question set; A standard answer module, configured to obtain a standard answer set corresponding to the test question set from an external database platform; A comparison and test module, configured to compare the consistency between the output answer set and the standard answer set through a large model, and obtain a test result of the question-and-answer intelligent agent according to the consistency comparison result.

13. A test intelligent agent for artificial intelligence, configured to execute the method according to any one of claims 1 to 11.

14. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-11.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

16. A computer program product, including a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-10.