A method, device, electronic device, storage medium and program for evaluating an intelligent agent

By determining the test case set by the agent role and using the agent evaluation model for quantitative evaluation, the problem that existing agent tests rely on manual evaluation is solved, and automated agent evaluation and efficient and accurate evaluation results are realized.

CN119621588BActive Publication Date: 2025-05-13BEIJING CHJ AUTOMOTIVE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510147019.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing agent test mainly relies on manual evaluation, which has problems such as large manpower investment, low evaluation efficiency and difficult to objectively quantify the evaluation results, resulting in inaccurate agent evaluation.

Method used

By determining the test case set of target agents based on the agent's role, generating test results, and quantitative evaluation is carried out based on the agent evaluation model, automatic evaluation and objective evaluation of agents are realized.

Benefits of technology

It improves the objectivity and unity of the evaluation of agents, ensures the accuracy and stability of the evaluation of agents, reduces manual intervention, and improves the evaluation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621588B_ABST
    Figure CN119621588B_ABST
Patent Text Reader

Abstract

The present invention discloses an agent evaluation method, device, electronic device, storage medium and program, wherein the method comprises: determining a test case set of a target agent based on an agent role; generating a test result of the target agent based on the test case set; and determining an evaluation result corresponding to each test result according to an agent evaluation model. The embodiment of the present invention can realize the automatic evaluation of the agent, perform quantitative evaluation on the agent evaluation process, and improve the objectivity of the agent evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technology, and in particular to an intelligent agent evaluation method, device, electronic device, storage medium and program. Background Art

[0002] An agent usually refers to an entity that can act autonomously in a certain environment, perceive environmental information, and make decisions based on its own goals, strategies, and capabilities to achieve specific tasks or goals. With the continuous development of deep learning, the integration of agents with the Internet of Things, big data, blockchain and other fields has deepened. The ability of agents in recognition, understanding, and generation has been rapidly improved, and agents are gradually playing a more important role in various fields. When an agent implements functions in different scenarios, it will play different roles. Due to the rich categories of agent roles and the obvious differences between agent roles, it is often necessary to test the performance of different agent roles when testing the agent. However, the current agent testing mainly relies on manual evaluation, which has the problems of large manpower investment and low evaluation efficiency. In addition, the evaluation process is mainly based on the subjective judgment of the evaluators, and the evaluation results are difficult to objectively quantify, resulting in inaccurate evaluation of the agent. Summary of the invention

[0003] The present invention provides an intelligent agent evaluation method, device, electronic device, storage medium and program to realize the automated evaluation of the intelligent agent, conduct quantitative evaluation on the intelligent agent evaluation process, and improve the objectivity of the intelligent agent evaluation.

[0004] According to one aspect of the present invention, a method for evaluating an intelligent agent is provided, wherein the method comprises:

[0005] Determine the test case set for the target agent based on the agent role;

[0006] Generate a test result of the target agent based on the test case set;

[0007] Determine the evaluation results corresponding to each of the test results according to the intelligent agent evaluation model.

[0008] According to another aspect of the present invention, there is provided an intelligent agent evaluation device, wherein the device comprises:

[0009] The evaluation rule module is used to determine the test case set of the target agent based on the agent role;

[0010] An evaluation execution module, used for generating a test result of the target agent based on the test case set;

[0011] The result evaluation module is used to determine the evaluation result corresponding to each of the test results according to the intelligent agent evaluation model.

[0012] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the intelligent agent evaluation method described in any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the intelligent agent evaluation method described in any embodiment of the present invention when executed.

[0017] According to another aspect of the present invention, a computer program product is provided, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the intelligent agent evaluation method as described in any embodiment of the present invention.

[0018] The technical solution of the embodiment of the present invention generates a test case set for the target intelligent agent based on the intelligent agent role, tests the target intelligent agent according to the test case set to obtain the test result of the target intelligent agent, evaluates the test result through an intelligent agent evaluation model to obtain the evaluation result of the target intelligent agent, determines the test scope of the target intelligent agent through the intelligent agent role, and tests the target intelligent agent based on the test case set corresponding to the test scope, thereby realizing automated testing of the target intelligent agent, and objectively evaluating the test results of the automated test through a unified intelligent agent evaluation model, which can improve the objectivity of the intelligent agent evaluation, ensure the uniformity of the intelligent agent evaluation, and help improve the stability of the intelligent agent operation.

[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 is a flow chart of an intelligent agent evaluation method provided according to Embodiment 1 of the present invention;

[0022] Figure 2 is a flow chart of another intelligent agent evaluation method provided according to the second embodiment of the present invention;

[0023] Figure 3 is a flow chart of another intelligent agent evaluation method provided according to Embodiment 3 of the present invention;

[0024] Figure 4 is a flow chart of another intelligent agent evaluation method provided according to Embodiment 4 of the present invention;

[0025] Figure 5 This is an example diagram of an intelligent agent evaluation method provided according to Embodiment 5 of the present invention;

[0026] Figure 6 is a structural diagram of an intelligent agent evaluation device provided according to Embodiment 6 of the present invention;

[0027] Figure 7 It is a schematic diagram of the structure of an electronic device for implementing the intelligent agent evaluation method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] Embodiment 1

[0031] Figure 11 is a flow chart of an intelligent agent evaluation method provided according to the first embodiment of the present invention. This embodiment is applicable to the case of intelligent agent evaluation. The method can be executed by an intelligent agent evaluation device. The device can be implemented in the form of hardware and / or software. The device can be configured in a server or terminal device. Figure 1 As shown, the method includes:

[0032] Step 110: Determine a test case set for the target agent based on the agent role.

[0033] Among them, the role of the agent can be the division of functions or tasks performed by different agents in a multi-agent system. The role of the agent can be divided into perceptual agents, decision-making agents, executive agents, coordination agents, and communication agents according to the tasks performed by the agents. The setting of the agent role can be determined based on the overall function and goal of the system. Taking the field of smart vehicles as an example, the roles of agents can include but are not limited to travel assistants, entertainment assistants, health management assistants, English teachers, encyclopedia teachers, and mental health teachers. A test case set can be a collection of measurement cases used to test a target agent. The test cases in the test case set can be one or several specific functional configurations for the target agent. The test cases can be composed of input data, preconditions, operation steps, and expected outputs. It can be understood that different test cases can be used to test different functions of the target agent.

[0034] In an embodiment of the present invention, the agent role and the target agent can be extracted. For example, the user can input the agent role and the target agent according to actual business needs, thereby determining the test case set for testing the target agent through the agent role. It can be understood that the agent role and the test case set can be stored in association, and the corresponding test case set can be found according to the identification information of the agent role, and the test case set can be used to test the target agent.

[0035] Step 120: Generate test results for the target agent based on the test case set.

[0036] Among them, the test result can be the result generated after each test case in the test case set is executed on the target executable body. The test result can include the conformity of the output result of the target executable body after processing according to the execution steps of the test case with the expected result in the test case. Each test case can include various corresponding test results.

[0037] In an embodiment of the present invention, the input data in the test case can be input into the target intelligent agent, so that the target intelligent agent is tested according to the execution steps in the test case, and the execution result output by the target agent can be compared with the output data of the test case to obtain the test result. It can be understood that the test process of the above-mentioned target agent based on the test case can be executed through automated evaluation rules, and the automated evaluation rules may include rules for applying the test case to the target agent. The automated evaluation rules may include but are not limited to evaluation rules based on visual perception technology, evaluation rules based on semantic understanding technology, evaluation rules based on generative evaluation technology, and evaluation rules based on scene recognition technology.

[0038] Exemplarily, after extracting the test case set, the automated evaluation rules can be extracted. The automated evaluation rules can be included in the test case set, or determined according to the type of test cases in the test case set. Exemplarily, the test cases in the test case set involve testing the image recognition function of the target intelligent body. According to the test objectives of the test cases, such as the image recognition function of the target intelligent body, the corresponding image recognition-based evaluation rules can be selected for the test objectives as the above-mentioned automated evaluation rules. It can be understood that each test case can have various corresponding automated evaluation rules. For example, if the test case needs to test the scene recognition function of the target executive body, the test rules based on scene recognition technology can be used as automated evaluation rules, while another test case involves functional testing of the semantic understanding of the target executive body, and the test rules based on semantic understanding can be used as automated evaluation rules.

[0039] Step 130: Determine the evaluation result corresponding to each test result according to the intelligent agent evaluation model.

[0040] Among them, the agent evaluation model can be an evaluation model constructed based on the rules of agent evaluation, the agent evaluation model can be composed of evaluation rules or evaluation indicators of at least one dimension, and the evaluation rules or evaluation indicators of each dimension in the agent evaluation model can have their own corresponding weight parameters. The agent evaluation model can output a final evaluation result based on all test results of the target agent, and the evaluation result can exist in the form of a quantitative score. The higher the value of the quantitative score, the better the target agent can be evaluated. Exemplarily, the agent evaluation model can at least include statistical rules for the test results of the agent, and the indicators generated by the statistical rules can be used as the evaluation results of the agent.

[0041] The embodiments of the present invention generate a test case set for the target agent based on the agent role, test the target agent according to the test case set to obtain corresponding test results, evaluate the test results through an agent evaluation model to obtain the evaluation results of the target agent, determine the test scope of the target agent through the agent role, and implement automated testing of the target agent based on the test cases of the corresponding test scope. Objectively evaluate the test results of the automated test through a unified agent evaluation model, which can improve the objectivity of the agent evaluation, ensure the uniformity of the agent evaluation, and help improve the stability of the agent operation.

[0042] Furthermore, on the basis of the above-mentioned embodiments of the invention, it also includes: extracting historical interaction behavior data of multiple users, and determining the similarity between the historical interaction behavior data and the preset intelligent agent roles in the intelligent agent role set; for the preset intelligent agent role, extracting the historical interaction behavior data with a similarity greater than a threshold as the target historical interaction behavior data; dividing the target historical interaction behavior data according to perception function, cognitive function, expression function and evolutionary function to obtain test cases corresponding to the preset intelligent agent role; saving each test case as a test case set for the preset intelligent agent role.

[0043] The historical interaction behavior data may be information about the use of agents of different agent roles by users, and may include behavior data input by users and feedback data generated by agents for the behavior data input by users. The historical interaction behavior data may include a label, which may correspond to the agent role, that is, the label may include information indicating the type of the agent role, and the label may be calibrated manually or in other ways. The historical interaction behavior data may include, but is not limited to, interaction information, vehicle information, and environmental information.

[0044] In an embodiment of the present invention, the similarity may be the degree of matching between the historical interaction behavior data and different preset intelligent agent roles in the intelligent agent role set. The similarity may include the degree of similarity between the historical interaction behavior data and the feature values ​​of the preset intelligent agent role or the degree of similarity between the historical interaction behavior data and the feature values ​​of the driving data associated with the preset intelligent agent role. The feature value may include but is not limited to word embedding features, semantic features, etc.

[0045] Specifically, historical interaction behavior data of multiple users and different preset intelligent agent roles pre-configured in the intelligent agent role set can be extracted, and similarity calculation can be performed between the historical interaction behavior data and the preset intelligent agent role. The similarity calculation can include calculating the similarity between the word embedding features of the historical driving data and the preset intelligent agent role, the similarity between the semantic features, etc., so as to determine the similarity between each historical interaction behavior data and different preset intelligent agent roles, and compare each similarity with a threshold value respectively. The threshold value can be a critical value for measuring the similarity between the historical interaction behavior data and the intelligent agent role. For each preset intelligent agent role, the historical interaction behavior data with a similarity greater than the threshold value can be used as the behavior data matching the preset intelligent agent role, and the matching The behavior data is recorded as the target historical interaction behavior data, and the target historical interaction behavior data can be divided according to the perception function, cognitive function, expression function and evolutionary function. Each different function can include the corresponding target historical interaction behavior data, and test cases can be generated for the target historical interaction behavior data of each function, so as to achieve a test case set that fully covers the role of the intelligent agent. The similarity threshold is used to ensure the degree of fit between the role of the intelligent agent and the real needs of the user, which can enhance the accuracy of the intelligent agent when testing based on the role of the intelligent agent. The matching target historical interaction behavior data is divided according to the perception function, cognitive function, expression function and evolutionary function, which can ensure that the generated test case set can cover all functions of the intelligent agent and enhance the comprehensiveness of the intelligent agent test. It can be understood that the above process of generating a test case set based on historical interaction behavior data can be generated by a deep learning network, or by other automated methods.

[0046] In some embodiments of the invention, the interaction information includes: interaction mode, interaction object and interaction scene. The interaction mode includes: single mode and multi-mode. Single mode includes: voice, touch, gesture, etc. The interaction object includes: vehicle computer function and intelligent entity function. The vehicle computer function includes: system function and vehicle adjustment function. The interaction scene includes: single-person scene and multi-person scene. Vehicle information includes: energy status and driving status. The energy status includes: power, fuel and charging status. The driving status includes: vehicle speed, gear and driving mode. Environmental information includes: time, weather, road and traffic conditions, etc.

[0047] In a specific example, the historical interaction behavior data is: [

[0049] [Voice + gesture, car window adjustment, commuting after get off work],

[0050] [Battery 30%, fuel 120km, low speed driving],

[0051] [Rainy day, 17 degrees Celsius, night, ring road, congestion],

[0052] ].

[0053] In this embodiment, the user historical interaction behavior data includes multiple types of data, and the test cases generated based on the multiple types of measured data can cover multiple interaction modes and environmental states, ensuring comprehensive scenario coverage.

[0054] In some embodiments of the invention, generating corresponding test cases according to the target historical interaction behavior data includes:

[0055] The target historical interaction behavior data and the knowledge base are input into the test case generation model to obtain the test cases corresponding to each target historical interaction behavior data; wherein the test case generation model is generated based on the training data set, and the training data set at least includes the interaction behavior data and the test cases corresponding to the interaction behavior data.

[0056] In an embodiment of the present invention, the test case generation model can be a pre-trained network model for generating test cases, the test case generation module can be implemented based on a recurrent neural network model or a long short-term memory model, the knowledge base can be interactive behavior data and test cases corresponding to the interactive behavior data, and the knowledge base can be generated by collecting historical test data of different users or simulated based on historical test data of some real users, etc. The test case generation model can be generated based on training of a training data set, and the training data set at least includes interactive behavior data and test cases corresponding to the interactive behavior data. The training data set can belong to the knowledge base. It is understandable that the knowledge base can be continuously updated over time. The knowledge base can assist in the model architecture design and parameter adjustment of the test case generation model, and expand the recognition mode and data processing trend of the test case generation model by updating the data in the knowledge base during the use of the model.

[0057] Specifically, a knowledge base and a test case generation model can be maintained. The test case generation model can be generated based on the training data set. The test case generation model can include but is not limited to a recurrent neural network model or a long short-term memory model. The knowledge base can continuously collect knowledge data related to the test case generation model during the use of the test case generation model. The knowledge data can include the user's interactive behavior data and test cases, and various evaluation indicators of the test case generation model, such as accuracy, recall rate, etc. It can be understood that the knowledge data in the knowledge base can be limited to the field of intelligent vehicles. Relevant knowledge in other related fields such as movies, the Internet, etc. can also be added to the knowledge base as knowledge data. When the target historical interactive behavior data is extracted, the target historical interactive behavior data and the knowledge base can be used as inputs to the test case generation model. The test case generation model can process the historical interactive behavior data and can rely on the knowledge correspondence relationship of the field-related knowledge corresponding to the test case in the knowledge base to determine the test case corresponding to the historical interactive behavior data.

[0058] Embodiment 2

[0059] Figure 2 is a flow chart of another agent evaluation method provided according to the second embodiment of the present invention. The embodiment of the present invention is a specific embodiment based on the above embodiment, and describes the process of determining the test case set. Figure 2 The method provided in the embodiment of the present invention specifically includes the following steps:

[0060] Step 210: extract the role indication information corresponding to the target intelligent agent.

[0061] The role indication information may be information indicating the agent role of the target agent being tested. The role indication information may indicate the target agent role by name or identification number. The role indication information may be input by a user when testing the target agent.

[0062] In an embodiment of the present invention, role indication information for a target intelligent agent can be received. For example, when a user tests a target intelligent agent, one or more target intelligent agent roles can be configured according to the system function or positioning of the target intelligent agent. The target intelligent agent role can be configured for the target intelligent agent by inputting role indication information in a visual interface.

[0063] Step 220: Search for the target agent role in the agent role set according to the role indication information.

[0064] Among them, the intelligent agent role set may include one or more intelligent agent role information configured for the intelligent agent. Different intelligent agent roles may correspond to different functions or task positioning of the intelligent agent. Taking the field of intelligent vehicles as an example, the intelligent agent roles may include but are not limited to travel assistants, entertainment assistants, health management assistants, English teachers, encyclopedia teachers, and mental health teachers, etc. Each intelligent map role may have various corresponding role indication information, and the role indication information may exist in the form of identification information or role name.

[0065] In an embodiment of the present invention, multiple preset intelligent agent roles can be pre-configured, and each preset intelligent agent role can have various corresponding role indication information, such as the role name of the intelligent agent role or the role number of the intelligent agent role, etc. When the role indication information is extracted, the preset intelligent agent role associated with the role indication information can be searched, and the found preset intelligent agent role can be used as the target intelligent agent role.

[0066] Step 230: extract the test cases stored in association with the target agent role to form a test case set.

[0067] Specifically, the target agent role can be associated with the test case, and different target agent roles can have corresponding relationships with their respective corresponding test cases. The corresponding relationship can be pre-configured, and the corresponding relationship can include storing all test cases corresponding to the target agent role in one location, or the target agent role and its corresponding test case have the same identification information. After determining the target agent role of the target agent, the corresponding test case can be extracted according to the target agent role, and the set of the above test cases can be used as the test case set of the target agent role.

[0068] Step 240: extract the input data, expected results, and test objects of each test case in the test case set.

[0069] Among them, the input data can be data used to input the target intelligent agent for testing. Based on the different types of test objects in the target intelligent agent, the input data can include images, voices, texts, videos, etc. The expected result can be the correct result expected to be output by the target intelligent agent. The test object can be the target function of the target intelligent agent to be tested. The function type of the test object can include at least one of perception function, cognitive function, expression function and evolutionary function. The perception function can include modal perception, situational perception, ecological perception, etc., while the cognitive function can include reasoning function and planning function, etc. The expression function can include calling function, content function, formal function, etc. The formal function can refer to the style and means of information output, while the evolutionary function can include learning function and memory function, etc.

[0070] In the embodiment of the present invention, a test case set of the target intelligent agent can be extracted, and for each test case in the test case set, input data, expected results, and test objects in the test case can be extracted.

[0071] Step 250: Search for a test execution script according to the functional type of the test object, wherein the functional type includes at least one of the following: perception function, cognitive function, expression function, and evolutionary function.

[0072] Among them, the function type can be the type of function to which the test object in the target intelligent body belongs, and the function type can include at least one of perception function, cognitive function, expression function and evolutionary function. The perception function can include modal perception, situational perception, ecological perception, etc., while the cognitive function can include reasoning function and planning function, etc. The expression function can include calling function, content function, formal function, etc. The formal function can refer to the style and means of information output, while the evolutionary function can include learning function and memory function, etc. The test execution script may be a test program used to test the test object. The test execution script may include large model processing rules, image recognition rules, generative evaluation rules, scene recognition rules, visual perception rules, semantic understanding rules, etc. There may be multiple test execution scripts, and the classification of test execution scripts may correspond to the functional type of the test control. For example, when the functional type of the test object is perception function, the corresponding test execution script may include visual perception rules and large model rules; when the functional type of the test object is authentication function, the corresponding test execution script may include semantic understanding rules and large model rules; when the functional type of the test object is expression function, the corresponding test execution script may include generative rules and image recognition rules; when the functional type of the test object is evolutionary function, the corresponding test execution script may include scene recognition rules and image recognition rules.

[0073] In the embodiment of the present invention, the functional type of the test object in the test case can be determined, and one or more corresponding test execution scripts can be selected according to the functional type.

[0074] Step 260: Call the test execution script to input the input data into the target intelligent agent, and compare the output result of the target intelligent agent with the expected result to obtain the test result.

[0075] In an embodiment of the present invention, the input data in the test case can be input into the target intelligent agent by calling the test execution script, so that the target intelligent agent executes according to the test steps specified in the test case. The test execution script can read the output result of the target intelligent agent and compare the output result with the expected result of the test case to obtain the test result. For example, the test result may include whether the output result is the same as or different from the expected result, and may also include the gap between the output result and the expected result, etc.

[0076] Exemplarily, taking the case where the test object includes cognitive functions, the input data of the test case may include route data to be planned, and the output data of the test case may include expected data corresponding to the input data, that is, route data generated by planning. The test execution script may input the route data to be planned into the intelligent agent, and the cognitive function of the intelligent agent processes the route data to be planned to obtain the planning output result. The test execution script calls the large model to compare the planning output result with the expected data to obtain the test result of the authentication function.

[0077] Step 270: Determine the evaluation result corresponding to each test result according to the intelligent agent evaluation model.

[0078] The embodiment of the present invention extracts role indication information for the target intelligent agent, finds the managed intelligent agent role as the target intelligent agent role according to the role indication information, queries the test case according to the target intelligent agent role to build a test case set, extracts the corresponding input data, expected results and test objects for each test case in the test case set, finds the corresponding test execution script according to the functional type of the test object, calls the test execution script to input the input data into the target intelligent agent to test the test object, and compares the output result generated by the target intelligent agent with the expected result through the test execution script to obtain the test result, and statistically analyzes the test result according to the intelligent agent evaluation model to obtain the evaluation result. The embodiment of the present invention determines the test case set of the target intelligent agent through the intelligent agent role, selects the pre-configured test execution script according to the functional type of the test case, and performs automatic testing on the target intelligent agent through the test execution script, which can improve the automation degree of the intelligent agent test, and provides a unified evaluation standard for the test result of the intelligent agent according to the intelligent agent evaluation model, which can improve the objectivity of the intelligent agent evaluation and help improve the stability of the intelligent agent operation.

[0079] Embodiment 3

[0080] Figure 3 is a flow chart of another agent evaluation method provided according to Embodiment 3 of the present invention. The embodiment of the present invention is a specific embodiment based on the above embodiment, and describes the evaluation generation process of the evaluation result. Figure 3 The method provided in the embodiment of the present invention specifically includes the following steps:

[0081] Step 310: Determine a test case set for the target agent based on the agent role.

[0082] Step 320: Generate test results for the target agent based on the test case set.

[0083] Step 330 extracts the functional evaluation dimensions within the intelligent agent evaluation model, wherein the functional evaluation dimensions include perception function, authentication function, expression function, and evolution function.

[0084] Among them, the functional evaluation dimension in the agent evaluation model may be a pre-configured dimension for evaluating the agent, the evaluation dimension may be divided based on the multimodal universal agent paradigm, and the evaluation dimension may at least include perception function, cognitive function, expression function, evolutionary function, etc. The evaluation index may be the index information for measuring the test result of the test case, and the evaluation index may include a quantitative rule for evaluating the test result.

[0085] In the embodiment of the present invention, the function evaluation dimensions such as perception function, authentication function, expression function and evolution function can be extracted from the pre-configured intelligent agent evaluation model, and different function evaluation dimensions can be used to evaluate the realization of the functions of the intelligent agent in different aspects. Furthermore, an evaluation index can be configured under each function evaluation dimension, and the evaluation index can be used to quantitatively analyze the realization of the functions of the intelligent agent, and the evaluation index can be pre-configured.

[0086] Step 340: Determine the target function evaluation dimension in each function evaluation dimension according to the function category to which the test result belongs.

[0087] In an embodiment of the present invention, the function category to which each test result belongs can be determined. The function category can be determined by the function category of the test object corresponding to the test result. For example, if the function category of the test object is cognitive function, the function category of the test result can also be cognitive function. Each test result can be mapped to different function evaluation dimensions in the intelligent agent evaluation model according to the function category to which it belongs. It can be understood that the function evaluation dimension includes perception function, authentication function, expression function and evolution function. The test result of the perception function is mapped to the function evaluation dimension of the perception function, so that the target function evaluation dimension is determined by the function category to which the test result belongs. Furthermore, the target function evaluation dimension can include one or more layers of evaluation indicators. Some or all of the evaluation indicators under the target function evaluation dimension corresponding to the test result can be used as the target evaluation indicator of the test result, and the test result can be quantitatively evaluated by the target evaluation indicator.

[0088] Step 350: extract the evaluation weights of each target function evaluation dimension in the intelligent agent evaluation model, and determine the evaluation results based on each evaluation weight and the corresponding test results.

[0089] The evaluation weight may be information reflecting the importance of each functional evaluation dimension in the agent evaluation process in the agent evaluation model. The larger the value of the evaluation weight, the more important the corresponding functional evaluation dimension is in evaluating the agent. The evaluation weight may be configured based on experiments or experience, or may be determined through a neural network model.

[0090] In an embodiment of the present invention, the evaluation weights configured for each target function evaluation dimension can be extracted from the intelligent agent evaluation model, and the evaluation weights can be applied to the corresponding evaluation results. This application can include determining the product or sum of the evaluation weights and the test results corresponding to each target evaluation indicator, and the results of the above-mentioned products and sums can be directly used as the evaluation results on the target function evaluation dimension.

[0091] According to an embodiment of the present invention, a test case set of a target agent is generated based on the agent role, test results of the target agent are generated according to the test case set, evaluation indicators of different functional evaluation dimensions in the agent evaluation model are extracted, and the test results are mapped to the target functional evaluation dimensions through different functional evaluation dimensions in the agent evaluation model according to the functional types to which they belong, and the indicator weights of each target functional evaluation dimension are extracted in the agent evaluation model, and the indicator weights are applied to the test results corresponding to the target functional evaluation dimensions, thereby obtaining evaluation results, determining the test scope of the target agent through the agent role, and automatically testing the target agent based on the test case set of the corresponding test scope, thereby realizing automated testing of the target agent, objectively evaluating the test results of the automated test through a unified agent evaluation model, thereby improving the objectivity of the agent evaluation, and adjusting the test results based on the importance of the evaluation impact of different functional evaluation dimensions, thereby improving the precision of the agent evaluation and ensuring the accuracy of the agent evaluation.

[0092] On the basis of the above-mentioned embodiments of the invention, the evaluation weights of each target function evaluation dimension in the intelligent agent evaluation model are extracted, and the evaluation results are determined based on each evaluation weight and the corresponding test results, including: extracting the evaluation weights of the target function evaluation dimensions of each test result in the intelligent agent evaluation model, wherein the sum of the evaluation weights of each test result is 1; for each test result, taking the product of the test result and the evaluation weight as the weighted test result; and taking the sum of each weighted test result as the evaluation result.

[0093] In an embodiment of the present invention, all test results can be generated by testing a test case set determined based on the role of the agent, and each test result can be obtained by testing at least one of the functions of the agent such as the perception function, cognitive function, expression function and evolutionary function of the test case set. In order to ensure the accuracy of the evaluation of the agent based on the role of the agent, the evaluation weight of the target function evaluation dimension corresponding to each test result can be extracted in the agent evaluation model, and the evaluation weight can reflect the importance of the target function evaluation dimension when evaluating the agent. When the test results are fully covered for the perception function, cognitive function, expression function and evolutionary function of the agent, the sum of the evaluation weights of each test result extracted in the agent evaluation model can be 1. The product between the test result and the corresponding evaluation weight can be determined for each test result, and the product result can be used as the weighted test result corresponding to the test result. The weighted test result can reflect the test result affected by different evaluation importances such as perception function, cognitive function, expression function and evolutionary function during the evaluation process. The sum of the weighted test results corresponding to all test results can be obtained, and the sum of the weighted test results corresponding to each test result can be used as the evaluation result of the agent evaluation. In an embodiment of the present invention, by obtaining test results covering all functions of the intelligent agent and adjusting the test results based on the importance of different functions in the evaluation process, an evaluation result that fully covers all functions of the intelligent agent can be obtained, thereby improving the integrity of the intelligent agent evaluation.

[0094] Further, based on the above-mentioned embodiments of the invention, the functional evaluation dimension includes at least one of the following: perceptual functional dimension, cognitive functional dimension, expressive functional dimension, and evolutionary functional dimension.

[0095] In an embodiment of the present invention, the perception function dimension may indicate the ability dimension of the intelligent agent to process data from one or more perception channels. The perception function dimension may be divided into multiple sub-dimensions, such as by different types of perception data, or by different stages of processing perception data. The cognitive function dimension may refer to the functional dimension of the intelligent agent to understand the perception and fusion information, discover the laws and relationships therein, and make inferences, predictions and decisions based on the laws and relationships. The authentication function dimension may be divided into one or more sub-function dimensions, and the division method may be divided according to the representation and storage method of knowledge, based on the types of cognitive reasoning and learning, etc. The representation function dimension may be a functional dimension that indicates the intelligent agent to generate one or more forms of expression according to the authentication results. The sub-dimension method of the representation function dimension may include but is not limited to division based on the form of expression, division based on the audience of expression, division based on the degree of personalization of expression, etc. The evolutionary function dimension may indicate the ability of the intelligent agent to continuously evolve in the functions of perception, cognition and expression. The evolutionary function dimension may be divided based on the evolutionary factors of perception, cognition and expression, so that the evolutionary function dimension may be composed of one or more sub-dimensions.

[0096] In some embodiments of the invention, it also includes: configuring the functional evaluation dimensions of the intelligent agent evaluation model according to the multimodal all-round intelligent agent paradigm; setting at least one layer of evaluation indicators in the intelligent agent evaluation model for each functional evaluation dimension; and determining the evaluation weights of the evaluation indicators of each functional evaluation dimension through an integrated learning algorithm.

[0097] In an embodiment of the present invention, a functional evaluation dimension can be constructed according to a multimodal all-round intelligent agent paradigm. The functional evaluation dimension can include but is not limited to perceptual function, cognitive function, expression function, evolutionary function, etc. Each functional evaluation dimension can be composed of multiple layers of evaluation indicators. Each layer of evaluation indicators can be a quantitative evaluation rule of a test result corresponding to a sub-dimension under the functional evaluation dimension. After configuring the evaluation indicators in each functional evaluation dimension, the evaluation weight corresponding to each evaluation indicator can be determined by an integrated learning algorithm. The integrated learning algorithm can include but is not limited to decision trees, neural networks, random forests, gradient boosting decision trees, etc. The embodiment of the present invention sets the functional evaluation dimension of the intelligent agent evaluation model through a multimodal all-round intelligent agent paradigm, which can ensure the reasonable setting of the functional evaluation dimension, ensure the integrity of the intelligent agent evaluation model, reduce the omission of the evaluation dimension of the intelligent agent, and improve the accuracy of the intelligent agent evaluation.

[0098] In some embodiments of the invention, determining the evaluation weight of the evaluation index of each functional evaluation dimension by an integrated learning algorithm includes:

[0099] The ensemble learning algorithm is called to determine the importance score of each evaluation dimension, wherein the ensemble learning algorithm is obtained by training based on evaluation indicator data with existing evaluation labels; each importance score is normalized to obtain the evaluation weight of each functional evaluation dimension.

[0100] The evaluation label may be information reflecting the importance of the evaluation index data in the evaluation process. The evaluation label may be determined based on historical information. The evaluation label may include a quantitative evaluation value and / or a classification label parameter.

[0101] In an embodiment of the present invention, an ensemble learning algorithm can be trained by using evaluation index data with evaluation labels. The training process can use evaluation index data with existing evaluation labels. After the ensemble learning algorithm is trained, the importance score values ​​of the functional evaluation dimensions are determined by the ensemble learning algorithm. The score values ​​corresponding to each functional evaluation dimension can be normalized so that the sum of the importance scores of the functional evaluation dimensions is 1. The score value of each functional evaluation dimension after normalization can be used as the evaluation weight. The embodiment of the present invention configures the evaluation weights of the functional evaluation dimensions through an ensemble learning algorithm, reduces the situation where the evaluation weights are set to abnormal values, improves the rationality of the evaluation weight setting, and can update the evaluation weights through the ensemble learning algorithm regularly or irregularly, which can enhance the flexibility of the intelligent agent evaluation model in conducting intelligent agent evaluation.

[0102] Furthermore, based on the above-mentioned embodiments of the invention, it also includes: generating an indicator evaluation report based on the evaluation results and the evaluation report template.

[0103] In an embodiment of the present invention, an evaluation report template can be extracted. There can be multiple types of evaluation report templates. Each evaluation result can be filled into the corresponding position in the evaluation report template according to its corresponding functional evaluation dimension. The evaluation report template filled with the evaluation results can be used as an indicator evaluation report.

[0104] Embodiment 4

[0105] Figure 4 is a flow chart of another agent evaluation method provided according to the fourth embodiment of the present invention. The embodiment of the present invention is a specific embodiment based on the above embodiment, and describes the process of generating the agent evaluation report. Figure 4 The method provided in the embodiment of the present invention specifically includes the following steps:

[0106] Step 410: Determine a test case set for the target agent based on the agent role.

[0107] Step 420: Generate test results for the target agent based on the test case set.

[0108] Step 430: Determine the evaluation result corresponding to each test result according to the intelligent agent evaluation model.

[0109] Step 440: Determine the missing data dimensions and data chart template in the evaluation report template.

[0110] Among them, the evaluation report template can be a template file for generating an evaluation report. The evaluation report template can be divided into multiple types according to different analysis purposes. The statistical analysis methods of the evaluation results in different evaluation report templates can be different. The evaluation report template can include an overall evaluation template for intelligent agents, a special analysis template for intelligent agents, a vertical development tracking template for intelligent agents, and a horizontal comparison template for similar intelligent agents.

[0111] In an embodiment of the present invention, the missing data dimension may be an agent function evaluation dimension that lacks actual data in the evaluation report template, and the missing data dimension may be set during the configuration of the evaluation report template. The data chart template may be a template for generating a chart based on the evaluation results, and the data chart template may at least include statistical analysis rules for the evaluation results and a display method for the chart.

[0112] Specifically, one or more evaluation report templates can be extracted, and the corresponding missing data dimensions and data chart templates can be extracted for each evaluation report template. It can be understood that the data chart template can be used to display statistical charts in the intelligent agent evaluation report corresponding to the evaluation report template. Each evaluation report template can have multiple corresponding data chart templates.

[0113] Step 450: When the functional evaluation dimension of the intelligent agent evaluation model matches the missing data dimension, the evaluation result of the evaluation dimension is filled into the vacant position corresponding to the missing data dimension in the evaluation report template.

[0114] In an embodiment of the present invention, the functional evaluation dimension can be matched with the missing data dimensions of different evaluation report templates respectively. If the functional evaluation dimension matches the missing data dimension, the match may include that the dimension name of the functional evaluation dimension is the same as that of the missing data dimension or the dimension identifier is the same, etc. Then, one or more evaluation results corresponding to the functional evaluation dimension can be extracted, and the above evaluation results can be filled into the vacant position corresponding to the missing data dimension matching the functional evaluation dimension in the evaluation report template, thereby realizing the filling of specific evaluation data in the evaluation report template.

[0115] Step 460: Count the evaluation results according to the indicators to be displayed in the data chart template to obtain evaluation statistical indicators.

[0116] The indicators to be displayed may be indicator values ​​generated according to a specific analysis and statistical method for the evaluation results, and the data chart template may include the specific analysis and statistical method for each indicator to be displayed.

[0117] In an embodiment of the present invention, an analysis and statistical method for the indicator to be displayed in the data chart template can be extracted, and one or more target evaluation results can be selected for processing according to the above analysis and statistical method. The statistical processing result can be used as an evaluation statistical indicator, which can be an indicator value corresponding to the indicator to be displayed.

[0118] Step 470: Generate an evaluation chart based on the evaluation statistical indicators according to the drawing method of the data chart template, and draw the evaluation chart into the intelligent agent evaluation report.

[0119] Specifically, the drawing method can be extracted from the data chart template according to the indicators to be displayed corresponding to the evaluation statistical indicators. The drawing method may include but is not limited to the chart display method and the chart display position, etc. The evaluation statistical indicators can be generated into an evaluation chart according to the corresponding drawing method, and the evaluation chart can be drawn in the intelligent agent evaluation report.

[0120] In the embodiment of the present invention, the test case set of the target agent is obtained through the agent role, and the test result of the target agent is generated according to the test case set, the test result is processed according to the agent evaluation model to obtain the evaluation result, the missing data dimension and the data chart template in the test report template are extracted, and the evaluation result is filled into the vacant position corresponding to the missing data dimension in the test report template according to the matching of the functional evaluation dimension of the evaluation result and the missing data dimension, and the evaluation result is analyzed and counted according to the indicators to be displayed in the data chart template to obtain the evaluation statistical indicators, and the evaluation statistical indicators are generated according to the drawing method of the data chart template, and the evaluation chart is drawn to the agent evaluation report. The embodiment of the present invention can realize the automatic generation of intelligent evaluation reports, display and analyze the test results corresponding to the test case set through the test report template, realize the standardization of agent evaluation, assist users to understand the evaluation results of the agent, facilitate the display of the chart content, improve the user's testing experience, reduce the sharing cost of the agent evaluation results, and assist users to coordinate the test adjustment of the agent.

[0121] Furthermore, based on the above-mentioned embodiments of the invention, the evaluation report template includes at least one of the following: an overall evaluation template for an intelligent agent, a special analysis template for an intelligent agent, a vertical development tracking template for an intelligent agent, and a horizontal comparison template for similar intelligent agents.

[0122] In an embodiment of the present invention, multiple evaluation report templates can be pre-configured for the purpose of evaluation of an intelligent agent. The evaluation report templates may include an overall evaluation template for an intelligent agent, a special analysis template for an intelligent agent, a vertical development tracking template for an intelligent agent, and a horizontal comparison template for similar intelligent agents, etc. Among them, the overall evaluation template for an intelligent agent may include a report template for overall analysis and statistics of all functional evaluation dimensions of an intelligent agent, and the intelligent agent evaluation template may display the evaluation results of each layer of evaluation indicators under each functional evaluation dimension respectively. The special analysis template for an intelligent agent may be a report template for analyzing and statistics of a specific functional dimension of an intelligent agent, and the special analysis template for an intelligent agent may be configured in multiple ways for different functional dimensions of an intelligent agent, and the special analysis template for an intelligent agent may include a chart template corresponding to a specific functional dimension and the analysis and statistics results of the evaluation results. The special analysis template for an intelligent agent may be a report template for analyzing and statistics of an intelligent agent according to a time sequence, and the special analysis template for an intelligent agent provides a comparative display channel for the evaluation results of the same evaluation indicator at at least two moments. The horizontal comparison template for similar intelligent agents may be a report template for analyzing and statistics of the evaluation results of different intelligent agents of the same type, and the horizontal comparison template for similar intelligent agents provides a comparative display channel for at least two intelligent agents to perform evaluation statistics for at least one functional dimension.

[0123] Specifically, the overall evaluation template of the intelligent agent, the special analysis template of the intelligent agent, the vertical development tracking template of the intelligent agent, and the horizontal comparison template of similar intelligent agents in the evaluation report template may respectively include at least one of a numerical display position and a chart display position. It can be understood that the evaluation report template may only include a data statistical report or a data chart.

[0124] In some embodiments of the invention, the evaluation report template includes an agent longitudinal development tracking template, and the agent evaluation report is generated based on the evaluation results and the evaluation report template, including:

[0125] Determine the missing data dimension of the intelligent agent longitudinal development tracking template, wherein the missing data dimension includes at least one of the functional evaluation dimension and the evaluation index; extract the evaluation results of at least two moments for the missing data dimension; fill the evaluation results of at least two moments into the vacant position corresponding to the missing data dimension in the intelligent agent longitudinal development tracking template to compare the evaluation results at different moments.

[0126] In an embodiment of the present invention, the evaluation report template may include an intelligent agent longitudinal development tracking template, which may reflect the performance changes of the intelligent agent in different functional evaluation dimensions or different evaluation indicators over time. The intelligent agent longitudinal development tracking report may select one or more functional evaluation dimensions of the intelligent agent evaluation model, or one or more evaluation indicators as missing data dimensions, and may select evaluation results of at least two moments of the corresponding functional evaluation dimensions or evaluation indicators according to the missing data dimensions. The above evaluation results may be filled into the vacant positions of the intelligent agent longitudinal development tracking template according to the corresponding missing data dimensions, so that the evaluation results of the corresponding missing data dimensions can be compared in the time dimension, thereby obtaining the performance changes of the intelligent agent over time.

[0127] Embodiment 5

[0128] Figure 5 is an example diagram of an agent evaluation method provided according to Embodiment 5 of the present invention, see Figure 5 The embodiment of the present invention provides an automated evaluation process of a multi-role agent based on a user experience perception model, which may include the following steps:

[0129] 1. Use case generation stage: Input a specific intelligent agent role, and automatically output a complete use case set through a pre-trained intelligent strategy algorithm module; the use case set includes a basic function use case set and an experience function use case set, among which the basic function can at least include recognition function, understanding function and reasoning function, while the experience function can at least include perception function, cognitive function, expression function and evolution function.

[0130] a. Among them, the intelligent strategy algorithm module is specifically used to: receive multiple test requirements, and combine existing test capabilities to automatically calculate and generate the optimal test strategy that takes into account both test coverage and test efficiency. The optimal test strategy includes the test scope and the evaluation object.

[0131] b. Pre-create a set of intelligent agent roles, including but not limited to: travel assistant, entertainment assistant, health management assistant, English teacher, encyclopedia teacher, mental health teacher, etc., and may also include role A generated according to user usage habits or user customization.

[0132] c. The test scope includes the complete set of test cases corresponding to the evaluation agent roles obtained through the agent role set.

[0133] 2. Automated evaluation stage: input specific evaluation use cases, select specific automated evaluation technology, implement the evaluation and output the evaluation results.

[0134] a. Obtain the user behavior sequence by analyzing and processing personnel, vehicle, and environmental data.

[0135] b. The user behavior sequence is converted into corresponding evaluation cases through the user behavior model.

[0136] 3. Result evaluation stage: The evaluation results of all test cases are input into the intelligent agent evaluation model to obtain the overall evaluation results of the product.

[0137] a. The agent evaluation model is divided into two parts: indicator system and indicator weight

[0138] b. The formation process of the indicator system is as follows:

[0139] i. The user experience perception model is based on the multimodal omnipotent agent paradigm and summarizes the four evaluation dimensions of perception-cognition-expression-iteration;

[0140] ii. Configure evaluation indicators of levels 2 to 5 under the four evaluation dimensions;

[0141] iii. Establish objective and quantitative standards based on the N evaluation indicators at the fifth level, which are also the screening criteria for the evaluation results.

[0142] c. The process of determining the indicator weight is as follows:

[0143] i. The indicator weight model is trained using an ensemble learning algorithm, which includes machine learning algorithms such as random forest, gradient boosting decision tree, and XGBoost;

[0144] ii. Use 70% of the existing evaluation index data as the training set and 30% as the test set; while using the training set to train the model, monitor the performance of the model on the test set;

[0145] iii. Obtain the importance ranking and scores of N evaluation indicator data based on the ensemble learning model;

[0146] iv. Normalize the N evaluation index data scores, compress them and sum them to 1 to obtain the evaluation index weights.

[0147] 4. Report output stage: Input the evaluation results into the automatic template generation tool to output the final evaluation report. Based on the proposed user experience perception model, an objective and quantifiable indicator system is formed.

[0148] In an embodiment of the present invention, the automatic template generation tool sets a variety of report templates, including but not limited to an overall evaluation template for an intelligent agent, a special analysis template for an intelligent agent, a vertical development tracking template for an intelligent agent, and a horizontal comparison template for similar intelligent agents. The automatic template generation tool can fill in data in the report template and draw charts with the overall evaluation results and evaluation results of the products generated in the above stages, thereby obtaining an evaluation report.

[0149] Embodiment 6

[0150] Figure 6 Schematic diagram of the structure of an intelligent agent evaluation device provided according to Embodiment 6 of the present invention. Figure 6 As shown, the device comprises:

[0151] The evaluation rule module 510 is used to determine the test case set of the target agent based on the agent role.

[0152] The evaluation execution module 520 is used to generate the test results of the target agent based on the test case set.

[0153] The result evaluation module 530 is used to determine the evaluation result corresponding to each of the test results according to the agent evaluation model.

[0154] In an embodiment of the present invention, an evaluation rule module generates a test case set for a target intelligent agent based on the role of the intelligent agent, an evaluation execution module generates a test result of the target intelligent agent according to the test case set, a result evaluation module evaluates the test result through an intelligent agent evaluation model to obtain an evaluation result of the target intelligent agent, determines the test scope of the target intelligent agent through the intelligent agent role, and implements automated testing of the target intelligent agent based on the test case set corresponding to the test scope. Objectively evaluating the test results of the automated test through a unified intelligent agent evaluation model can improve the objectivity of the intelligent agent evaluation, ensure the uniformity of the intelligent agent evaluation, and help improve the stability of the intelligent agent operation.

[0155] Based on the above-mentioned embodiment of the invention, the evaluation rule module 510 is specifically used to: extract role indication information corresponding to the target intelligent agent; search for the target intelligent agent role in the intelligent agent role set according to the role indication information; extract the test cases stored in association with the target intelligent agent role to form the test case set.

[0156] Based on the above embodiment of the invention, a test case set module is also included, and the test case set module includes:

[0157] The similarity unit is used to extract historical interaction behavior data of multiple users and determine the similarity between the historical interaction behavior data and preset intelligent agent roles in the intelligent agent role set.

[0158] A use case generation unit is used to extract the historical interaction behavior data with a similarity greater than a threshold as target historical interaction behavior data for the preset intelligent agent role; divide the target historical interaction behavior data according to perception function, cognitive function, expression function and evolutionary function to obtain a test case corresponding to the preset intelligent agent role.

[0159] A use case set unit is used to save each of the test cases as the test case set of the preset intelligent agent role.

[0160] Based on the above-mentioned embodiment of the invention, the evaluation execution module 520 includes: extracting the input data, expected results and test objects of each test case in the test case set; searching for the test execution script according to the functional type of the test object; calling the test execution script to input the input data into the target intelligent agent, and comparing the output result of the target intelligent agent with the expected result to obtain the test result; wherein the functional type includes at least one of the following: perception function, cognitive function, expression function and evolutionary function.

[0161] Based on the above-mentioned embodiment of the invention, the result evaluation module 530 includes:

[0162] The indicator extraction unit is used to extract the functional evaluation dimensions within the intelligent agent evaluation model, wherein the functional evaluation dimensions include perception function, authentication function, expression function and evolution function.

[0163] The target dimension unit is used to determine the target function evaluation dimension in each function evaluation dimension according to the function category to which the test result belongs.

[0164] A weight processing unit is used to extract the evaluation weight of the target function evaluation dimension in the intelligent agent evaluation model, and determine the evaluation result based on the evaluation weight and the test result.

[0165] Based on the above-mentioned embodiment of the invention, the weight processing unit extracts the evaluation weight of each target function evaluation dimension in the agent evaluation model, and determines the evaluation result based on each evaluation weight and the corresponding test result, including:

[0166] The evaluation weight of the target function evaluation dimension of each test result is extracted in the agent evaluation model, wherein the sum of the evaluation weights of the test results is 1; for each test result, the product of the test result and the evaluation weight is taken as the weighted test result; and the sum of the weighted test results is taken as the evaluation result.

[0167] Based on the above-mentioned embodiments of the invention, the functional evaluation dimension includes at least one of the following: perceptual functional dimension, cognitive functional dimension, expressive functional dimension, and evolutionary functional dimension.

[0168] On the basis of the above-mentioned embodiments of the invention, it also includes: a parameter determination module, which is used to configure the functional evaluation dimensions of the intelligent agent evaluation model according to the multimodal all-round intelligent agent paradigm; set at least one layer of evaluation indicators in the intelligent agent evaluation model for each functional evaluation dimension; and determine the evaluation weights of the evaluation indicators of each functional evaluation dimension through an integrated learning algorithm.

[0169] Based on the above-mentioned embodiment of the invention, the evaluation weights of the evaluation indicators of each function evaluation dimension are determined by an integrated learning algorithm, including:

[0170] Calling the ensemble learning algorithm to determine the importance score of each of the function evaluation dimensions, wherein the ensemble learning algorithm is trained based on evaluation indicator data with existing evaluation labels;

[0171] Each of the importance scores is normalized to obtain the evaluation weight of each of the function evaluation dimensions.

[0172] Based on the above-mentioned embodiment of the invention, it also includes: a report generation module, which is used to generate an intelligent agent evaluation report based on the evaluation results and the evaluation report template.

[0173] Based on the above-mentioned embodiment of the invention, the report generation module includes:

[0174] The template parsing unit is used to determine the missing data dimensions and data chart template in the evaluation report template, wherein the evaluation report template includes at least one of the following: an overall evaluation template for an intelligent agent, a special analysis template for an intelligent agent, a vertical development tracking template for an intelligent agent, and a horizontal comparison template for similar intelligent agents.

[0175] A data filling unit, configured to fill the evaluation result of the function evaluation dimension into the vacant position corresponding to the missing data dimension in the evaluation report template when the function evaluation dimension matches the missing data dimension;

[0176] A data statistics unit, used for collecting statistics of each evaluation result according to the indicators to be displayed in the data chart template to obtain evaluation statistical indicators;

[0177] A chart drawing unit is used to generate an evaluation chart based on the evaluation statistical indicators according to the drawing method of the data chart template, and draw the evaluation chart into the intelligent agent evaluation report.

[0178] Based on the above-mentioned embodiments of the invention, the report generation module is specifically used to: determine the missing data dimension of the intelligent agent longitudinal development tracking template, wherein the missing data dimension includes at least one of the functional evaluation dimension and the evaluation index; for the missing data dimension, extract the evaluation results of at least two moments; fill the evaluation results of the at least two moments into the vacant position corresponding to the missing data dimension in the intelligent agent longitudinal development tracking template to compare the evaluation results at different moments.

[0179] The intelligent agent evaluation device provided in the embodiment of the present invention can execute the intelligent agent evaluation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0180] Embodiment 7

[0181] Figure 7 : is a schematic diagram of the structure of an electronic device that implements the intelligent agent evaluation method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0182] like Figure 7 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0183] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0184] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the intelligent agent evaluation method.

[0185] In some embodiments, the agent evaluation method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the agent evaluation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the agent evaluation method in any other appropriate manner (e.g., by means of firmware).

[0186] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0187] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0188] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0190] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0191] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0192] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0193] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An agent evaluation method, characterized in that: The method comprises: Extracting historical interaction behavior data of multiple users, and determining the similarity between the historical interaction behavior data and a preset intelligent agent role in the intelligent agent role set; for the preset intelligent agent role, extracting the historical interaction behavior data with the similarity greater than a threshold as target historical interaction behavior data; dividing the target historical interaction behavior data according to perception function, cognitive function, expression function and evolutionary function to obtain test cases corresponding to the preset intelligent agent role; saving each of the test cases as a test case set for the preset intelligent agent role; Determine the test case set for the target agent based on the agent role; Generate a test result of the target agent based on the test case set; Determining the evaluation results corresponding to each of the test results according to the intelligent agent evaluation model, including: extracting the functional evaluation dimensions within the intelligent agent evaluation model, wherein the functional evaluation dimensions include perceptual function, cognitive function, expressive function and evolutionary function; determining the target functional evaluation dimensions in each of the functional evaluation dimensions according to the functional category to which the test results belong; extracting the evaluation weights of each of the target functional evaluation dimensions in the intelligent agent evaluation model, and determining the evaluation results based on each of the evaluation weights and the corresponding test results.

2. The method according to claim 1, characterized in that: The test case set for determining the target agent based on the agent role includes: Extracting role indication information corresponding to the target agent; Searching for a target agent role in the agent role set according to the role indication information; The test cases stored in association with the target agent role are extracted to form the test case set.

3. The method according to any one of claims 1 or 2, characterized in that: The generating the test result of the target agent based on the test case set comprises: Extracting input data, expected results, and test objects of each test case in the test case set; Searching for a test execution script according to the functional type of the test object; Calling the test execution script to input the input data into the target intelligent agent, and comparing the output result of the target intelligent agent with the expected result to obtain the test result; Among them, the functional types include at least one of the following: perception function, cognitive function, expression function and evolutionary function.

4. The method according to claim 1, characterized in that: The extracting the evaluation weight of each of the target function evaluation dimensions in the agent evaluation model, and determining the evaluation result based on each of the evaluation weights and the corresponding test result, includes: Extracting the evaluation weight of the target function evaluation dimension of each of the test results in the agent evaluation model, wherein the sum of the evaluation weights of the test results is 1; For each of the test results, taking the product of the test result and the evaluation weight as a weighted test result; The sum of the weighted test results is taken as the evaluation result.

5. The method according to claim 1, characterized in that: Also includes: Configuring the functional evaluation dimensions of the agent evaluation model according to a multimodal all-round agent paradigm; For each of the functional evaluation dimensions, at least one layer of evaluation indicators is set in the agent evaluation model; The evaluation weights of the evaluation indicators of each functional evaluation dimension are determined by an integrated learning algorithm.

6. The method according to claim 5, characterized in that: Determining the evaluation weight of the evaluation index of each functional evaluation dimension by an integrated learning algorithm includes: Calling the ensemble learning algorithm to determine the importance score of each of the function evaluation dimensions, wherein the ensemble learning algorithm is trained based on evaluation indicator data with existing evaluation labels; Each of the importance scores is normalized to obtain the evaluation weight of each of the function evaluation dimensions.

7. The method according to claim 1, characterized in that: Also includes: Generating an agent evaluation report based on the evaluation results and the evaluation report template, wherein generating the agent evaluation report based on the evaluation results and the evaluation report template comprises: Determine the missing data dimensions and data chart templates in the evaluation report template, wherein the evaluation report template includes at least one of the following: an agent overall evaluation template, an agent thematic analysis template, an agent vertical development tracking template, and a similar agent horizontal comparison template; When the function evaluation dimension of the agent evaluation model matches the missing data dimension, filling the evaluation result of the function evaluation dimension into the vacant position corresponding to the missing data dimension in the evaluation report template; According to the indicators to be displayed in the data chart template, statistics are collected on each of the evaluation results to obtain evaluation statistical indicators; The evaluation statistical indicators are generated into an evaluation chart according to the drawing method of the data chart template, and the evaluation chart is drawn into the intelligent agent evaluation report.

8. The method according to claim 7, characterized in that: The evaluation report template includes an agent vertical development tracking template, and the agent evaluation report is generated based on the evaluation result and the evaluation report template, including: Determining the missing data dimension of the agent longitudinal development tracking template, wherein the missing data dimension includes at least one of a functional evaluation dimension and an evaluation index; For the missing data dimension, extract the evaluation results at at least two moments; The evaluation results of the at least two moments are filled into the vacant positions corresponding to the missing data dimensions in the intelligent agent longitudinal development tracking template to compare the evaluation results at different moments.

9. An intelligent agent evaluation device, characterized in that: The device comprises: A test case set module is used to extract historical interaction behavior data of multiple users, and determine the similarity between the historical interaction behavior data and a preset intelligent agent role in the intelligent agent role set; for the preset intelligent agent role, extract the historical interaction behavior data with the similarity greater than a threshold as the target historical interaction behavior data; divide the target historical interaction behavior data according to perception function, cognitive function, expression function and evolutionary function to obtain the test case corresponding to the preset intelligent agent role; save each of the test cases as a test case set for the preset intelligent agent role; The evaluation rule module is used to determine the test case set of the target agent based on the agent role; An evaluation execution module, used for generating a test result of the target agent based on the test case set; A result evaluation module is used to determine the evaluation results corresponding to each of the test results according to the intelligent agent evaluation model, including: extracting the functional evaluation dimensions within the intelligent agent evaluation model, wherein the functional evaluation dimensions include perceptual function, cognitive function, expressive function and evolutionary function; determining the target functional evaluation dimensions in each of the functional evaluation dimensions according to the functional category to which the test results belong; extracting the evaluation weights of each of the target functional evaluation dimensions in the intelligent agent evaluation model, and determining the evaluation results based on each of the evaluation weights and the corresponding test results.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the intelligent agent evaluation method described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the intelligent agent evaluation method described in any one of claims 1-8 when executed.

12. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the intelligent agent evaluation method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Large model agent evaluation method and device

    CN118747211A