Intelligent agent evaluation method and device, equipment, readable storage medium and product
By combining the real needs of the target user with randomly generated scenario variables, the input content of the intelligent agent is updated, which solves the problem of low accuracy of intelligent agent evaluation in dynamic scenarios and realizes a comprehensive evaluation of its processing capabilities.
Patent Information
- Application Number
- CN202511338131.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-01-13
AI Technical Summary
In existing technologies, agent evaluation methods cannot accurately reflect their processing capabilities in dynamically changing scenarios, resulting in a significant discrepancy between the evaluation results and the actual effects.
By using the target user's current input content and demand tag system, a demand identification model is used to determine demand description information, randomly generate scenario variables, update the input content, and process it through an intelligent agent to obtain evaluation results, simulating unexpected emergencies and improving evaluation accuracy.
It enables a comprehensive and accurate assessment of the agent's processing capabilities in dynamic scenarios, improving the accuracy and effectiveness of the assessment.
Smart Images

Figure CN121328609A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for evaluating intelligent agents. Background Technology
[0002] With the development of computer technology, artificial intelligence technology has advanced rapidly. Intelligent agents, as core carriers with autonomous decision-making and environmental interaction capabilities, have been widely applied in complex scenarios such as customer service, autonomous driving, medical assistance, and industrial control. As the tasks undertaken by intelligent agents become increasingly critical, evaluating their processing capabilities has become a core challenge for promoting technology implementation, ensuring system reliability, and fostering the healthy development of the industry.
[0003] In related technologies, agents are typically evaluated using static evaluation sets with fixed scenarios. However, in the real world, scenarios are dynamic and changing. Therefore, using static evaluation sets for evaluation can lead to a large discrepancy between the evaluation results and the actual effects, resulting in low accuracy in agent evaluation. Summary of the Invention
[0004] Therefore, it is necessary to provide an intelligent agent evaluation method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the accuracy of intelligent agent evaluation in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for evaluating an intelligent agent, including:
[0006] Based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags.
[0007] Based on the current input content and the requirement description information, a scene variable for the target scene is randomly generated. The current input content is then updated based on the scene variable to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variable to the target scene.
[0008] Based on the updated input content, the intelligent agent processes the data to obtain a processing result. Based on the processing result, the intelligent agent is evaluated to obtain an evaluation result.
[0009] Secondly, this application also provides an evaluation device for an intelligent agent, comprising:
[0010] The information acquisition module is used to determine the target user's demand description information based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags.
[0011] The content update module is used to randomly generate scene variables for the target scene based on the current input content and the requirement description information, and update the current input content based on the scene variables to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variables to the target scene.
[0012] The agent evaluation module is used to process the updated input content through an agent to obtain a processing result, and to evaluate the agent based on the processing result to obtain an evaluation result.
[0013] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0014] Based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags.
[0015] Based on the current input content and the requirement description information, a scene variable for the target scene is randomly generated. The current input content is then updated based on the scene variable to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variable to the target scene.
[0016] Based on the updated input content, the intelligent agent processes the data to obtain a processing result. Based on the processing result, the intelligent agent is evaluated to obtain an evaluation result.
[0017] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0018] Based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags.
[0019] Based on the current input content and the requirement description information, a scene variable for the target scene is randomly generated. The current input content is then updated based on the scene variable to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variable to the target scene.
[0020] Based on the updated input content, the intelligent agent processes the data to obtain a processing result. Based on the processing result, the intelligent agent is evaluated to obtain an evaluation result.
[0021] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0022] Based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags.
[0023] Based on the current input content and the requirement description information, a scene variable for the target scene is randomly generated. The current input content is then updated based on the scene variable to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variable to the target scene.
[0024] Based on the updated input content, the intelligent agent processes the data to obtain a processing result. Based on the processing result, the intelligent agent is evaluated to obtain an evaluation result.
[0025] The aforementioned evaluation method, apparatus, computer equipment, computer-readable storage medium, and computer program product for intelligent agents, through a demand identification model, accurately identifies the target user's true needs based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content. This demand identification model determines the target user's demand description information, which includes target demand tags and quantitative indicators used to quantify these tags. Based on the current input content and demand description information, scenario variables for the target scenario are randomly generated. In other words, by combining the target user's true needs, unexpected unforeseen situations that may occur in the target scenario are randomly generated. Therefore, the current input content is updated based on the scenario variables, resulting in updated input content. The scenario involved in the updated input content is the target scenario obtained by adding scenario variables. Thus, the updated input content is based on the current input content with the addition of randomly generated scenario variables. Based on this updated input content, the intelligent agent processes the data to obtain a processing result. Based on the processing result, the intelligent agent is evaluated to obtain an evaluation result. This allows for a comprehensive and accurate evaluation of the intelligent agent's processing capability under the dynamic addition of unexpected scenario variables to the target scenario, effectively improving the accuracy of the intelligent agent evaluation. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a diagram illustrating the application environment of an agent evaluation method in one embodiment.
[0028] Figure 2 This is a flowchart illustrating an agent evaluation method in one embodiment;
[0029] Figure 3 This is a schematic diagram of the steps for determining the requirement description information in one embodiment;
[0030] Figure 4 This is a schematic diagram illustrating the updating of input content in one embodiment;
[0031] Figure 5 This is a schematic diagram of the evaluation process in one embodiment;
[0032] Figure 6 This is a schematic diagram of the feedback process in one embodiment;
[0033] Figure 7This is a schematic diagram of the agent evaluation process in one embodiment;
[0034] Figure 8 This is a structural block diagram of an evaluation device for an agent in one embodiment;
[0035] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0037] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0038] In introducing the embodiments of this application, the concept of an intelligent agent will first be explained:
[0039] An intelligent agent is a proxy capable of perceiving its environment and taking actions to achieve specific goals. It can be software, hardware, or a system, possessing autonomy, adaptability, and interactivity. Intelligent agents perceive changes in the environment (e.g., through sensors or data input), make judgments and decisions based on their learned knowledge and algorithms, and then execute actions to influence the environment or achieve predetermined goals. Intelligent agents are widely used in the field of artificial intelligence, commonly found in automated systems, robots, virtual assistants, and game characters. Their core strength lies in their ability to learn autonomously and continuously evolve to better complete tasks and adapt to complex environments.
[0040] The agent evaluation method provided in this application can be applied to, for example... Figure 1 In the application environment shown, computer device 102 communicates with agent 104 via a network.
[0041] In some embodiments, computer device 102 acquires the current input content of a target user, and based on the current input content and the demand tag system corresponding to the target scenario involved in the current input content, determines the demand description information of the target user through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags. Based on the current input content and the demand description information, a scenario variable for the target scenario is randomly generated. The current input content is updated based on the scenario variable to obtain updated input content. The scenario involved in the updated input content is the scenario obtained by adding the scenario variable to the target scenario. Computer device 102 sends the updated input content to intelligent agent 104, which processes the updated input content to obtain a processing result. Intelligent agent 104 returns the processing result to computer device 102, and computer device 102 evaluates the intelligent agent based on the processing result to obtain an evaluation result.
[0042] The computer device 102 can be a terminal or a server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0043] In one exemplary embodiment, such as Figure 2 As shown, an evaluation method for an intelligent agent is provided, which can be applied to... Figure 1 Taking computer device 102 as an example, the explanation includes the following steps 202 to 206. Wherein:
[0044] Step 202: Based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through the demand identification model. The demand description information includes the target demand tags of the target user and the quantitative indicators used to quantify the target demand tags.
[0045] Here, the target user's current input content refers to the content newly entered by the target user at the current moment. The content format of the current input content can be video, text, images, etc. Optionally, the current input content is obtained by the target user through the content input box on the target application. For example, in response to the road condition video sending operation in the intelligent driving application, the computer device uses the current road condition video of the target user's vehicle as the current input content. This current road condition video shows the situation ahead of the vehicle, such as the traffic light situation on a sunny road. For example, in response to the content sending operation in the conversation application, the content entered by the target user in the conversation application is used as the current input content. For example, if the target user initiates a query in the conversation application, such as "When can this solution be implemented?", the current input content is used.
[0046] The target scenario involved in the current input content can be the scenario indicated in the current input content, or a scenario associated with the current input content. For example, if the current input content is a video of current road conditions, the corresponding target scenario is an intelligent driving scenario. Or, if the current input content is "When can this solution be implemented?", the corresponding target scenario is a project management scenario. In some embodiments, the current input content may contain keywords indicating the target scenario, such as the keyword "intelligent driving," thereby determining the target scenario. In other embodiments, if the current input content does not contain keywords indicating the target scenario, it can be determined based on the application from which the current input content originates. For example, if it is sent through an intelligent driving application, the corresponding target scenario is an intelligent driving scenario. The target scenario can be a scenario in multiple fields such as hot topic prediction, content alerts, risk management, script recommendations, business negotiations, medical diagnosis, and educational tutoring.
[0047] Each scenario has a corresponding demand tagging system, which includes demand tags specific to that scenario. For example, in the intelligent driving scenario, demand tags could be safety needs and timeliness needs; in the project management scenario, demand tags could be project completion needs and time cost needs. The demand identification model is used to identify the demand for the current input content, which can be implicit or implicit. In other words, the demand identification model captures the target user's actual needs for the current input content. Target demand tags refer to the demand tags identified by the demand identification model that match the current input content within the demand tagging system corresponding to the target scenario. Target demand tags reflect the target user's true user needs. Quantitative indicators are used to evaluate whether the target demand tags are completed or satisfied while meeting the rules. Quantitative indicators include the objectives and rules for the target demand tags. Rules can be the accuracy of demand identification for the corresponding demand tags or the completeness of the response to potential demands, etc. The corresponding objective can be that the agent processes the user's input content and provides a processing result, that is, the agent completes the processing. Of course, corresponding goals and rules can also be set for specific scenarios. For example, the goal is for the agent to fully process the input content, that is, for the agent to give a processing result about the input content. The corresponding rules can be the conditions that need to be followed in the scenario involving the input content. For example, in the intelligent driving scenario, the target requirement label is driving safety, the corresponding goal is safe passage, and the corresponding rules are red light stop, green light go, yellow light, etc.
[0048] Optionally, the computer device acquires the current input content of the target user and identifies the target scenario involved in the current input content by performing at least one of scenario keyword recognition and content understanding. The computer device filters out a demand tag system corresponding to the target scenario from multiple candidate demand systems. Based on the current input content and the involved demand tag system, the computer device determines the target user's target demand tags and quantitative indicators used to quantify the target demand tags through a demand identification model to obtain demand description information.
[0049] In some embodiments, the computer device performs scene keyword recognition on the current input content to obtain a recognition result, and then performs content understanding on the current input content using a content understanding model to obtain an understanding result. If both the recognition result and the understanding result detect scene information, and the detected scene information indicates the same scene, the indicated scene is taken as the target scene.
[0050] In some embodiments, to facilitate the determination of the target scenario, the method further includes: after obtaining the content input by the target user, if the content does not contain scenario information, the computer device can confirm the scenario corresponding to the content with the user through polling; after determining the scenario confirmed by the user, the scenario information of the confirmed scenario is added to the content to obtain the currently input content. This ensures that subsequent scenario information can be directly obtained through keyword retrieval without further content understanding, thus improving the efficiency of scenario confirmation.
[0051] Step 204: Based on the current input content and requirement description information, randomly generate scene variables for the target scenario, update the current input content based on the scene variables, and obtain updated input content. The scenario involved in the updated input content is the scenario obtained by adding scene variables to the target scenario.
[0052] Scenario variables are variables used to change the external conditions of the target scenario. For example, scenario variables can increase or decrease the severity of the target scenario. In the context of intelligent driving, scenario variables could be weather, such as severe weather (wind, rain, etc.) or mild weather (sunny days, etc.). In business negotiation scenarios, corresponding scenario variables could be market fluctuations (such as the range of product price fluctuations) or competitors suddenly proposing new terms. In medical diagnosis scenarios, corresponding scenario variables could be sudden complications in patients or temporary malfunctions in diagnostic equipment.
[0053] In some embodiments, the scenario variables are unexpected scenario variables, such as sudden interference, target drift, sudden changes in environmental parameters, etc.
[0054] Updating the current input based on scene variables essentially involves adding scene variables to the target scene described by the current input, resulting in the input content of the target scene after the addition. For example, in an autonomous driving scenario, if the current input is a video of the current road conditions (i.e., the video shows traffic lights on a sunny road), and the scene variable is "rainy," then the updated input would be a video showing traffic lights on a rainy road.
[0055] Optionally, the computer device generates scene variables for the target scene based on the current input content and the requirement description information through a random variable generation model. This random variable generation model is a model built on a neural network, or it can be a large language model.
[0056] For example, a computer device acquires scene attribute information of a target scene, and based on the scene attribute information, current input content, and requirement description information, uses a random variable generation algorithm to determine the variable type and value range of scene variables, and then determines the scene variables based on the variable type and value range. The scene attribute information can be information reflecting the characteristics of the scene.
[0057] Step 206: Based on the updated input content, the intelligent agent processes the data to obtain the processing result. Based on the processing result, the intelligent agent is evaluated to obtain the evaluation result.
[0058] The processing result is the outcome obtained by the agent after analyzing and processing the updated input content. This result includes the inference result, which is the response to the updated input content. For example, if the updated input content is a video showing the traffic light situation on a road in rainy weather, the corresponding inference result would be "proceed slowly".
[0059] The evaluation result is an assessment of the agent's processing capabilities, that is, based on the processing result, it evaluates the agent's perception and adaptation capabilities in the face of unexpected scene variables, and whether it can make accurate inferences.
[0060] Optionally, the computer device inputs updated input content into the agent, which then analyzes and reasons about the updated input content to obtain a processing result. Based on the processing result, the computer device scores the result from the perception, reasoning, and execution dimensions. Based on the scores for each dimension, the final score of the agent is determined, and an evaluation result is obtained. If the final score is lower than a preset threshold, the evaluation result indicates that the agent has failed and needs adjustment based on external feedback. If the final score is higher than or equal to the preset threshold, the evaluation result indicates that the agent has passed. The perception dimension evaluates the agent's ability to perceive the updated input content, i.e., whether it correctly perceives the requirement. The reasoning dimension evaluates the agent's ability to reason about the input content, i.e., whether it can correctly respond to the input content. The execution dimension evaluates whether the agent's recognized requirement matches the derived reasoning result.
[0061] In the aforementioned evaluation method for intelligent agents, a demand identification model is used to accurately identify the target user's true needs based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content. This demand identification model determines the target user's demand description information, which includes target demand tags and quantitative indicators used to quantify these tags. Based on the current input content and demand description information, scenario variables for the target scenario are randomly generated. In other words, by combining the target user's true needs, unexpected unforeseen situations that may occur in the target scenario are randomly generated. Therefore, the current input content is updated based on these scenario variables, resulting in updated input content. The scenario involved in the updated input content is the target scenario with the added scenario variables. Thus, the updated input content is based on the current input content with the addition of randomly generated scenario variables. Based on this updated input content, the intelligent agent processes the data to obtain a processing result. Based on this processing result, the intelligent agent is evaluated to obtain an evaluation result. This method comprehensively and accurately evaluates the intelligent agent's processing capability under the dynamic addition of unexpected scenario variables to the target scenario, effectively improving the accuracy of the intelligent agent evaluation.
[0062] In some embodiments, based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through a demand identification model, including: obtaining a historical sample set, where each historical sample in the historical sample set includes historical input content; processing the historical input content through an intelligent agent to obtain historical processing results; using an emotion recognition model to perform emotion recognition on the current input content to obtain emotion recognition results; and determining the target user's demand description information through a demand identification model based on the historical sample set, the emotion recognition results, the current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario.
[0063] The historical sample set includes multiple historical samples. The historical input content in these samples may or may not be input by the target user; there are no specific limitations. The sentiment recognition model identifies the sentiment information of the current input content to determine whether the target user is in a state of satisfaction, dissatisfaction, anxiety, or neutrality. This sentiment recognition result indicates the target user's emotional state and can also reflect implicit needs. Each scenario has a corresponding demand labeling system, which includes the needs appropriate for that scenario. For example, in an e-commerce shopping scenario, implicit needs in the corresponding demand labeling system can be categorized into concerns about product quality, pursuit of cost-effectiveness, and expectations of personalized recommendations. In an intelligent customer service scenario, demand labels in the corresponding demand labeling system could include requirements for problem-solving speed, expectations of professional solutions, and needs for emotional care. For example, the demand labeling system can be built based on historical experience or by summarizing a complete set of needs. Based on knowledge of a specific domain, such as knowledge of the corresponding scenario, the needs in the demand set are categorized to obtain the needs matching the corresponding scenario, thus obtaining the corresponding demand labeling system.
[0064] Optionally, after acquiring a historical sample set, the computer device uses an emotion recognition model to perform emotion recognition on the current input content, obtaining the emotion tendency and intensity, thus yielding the emotion recognition result. This emotion recognition model can be a model built using machine learning classification algorithms, or it can utilize an emotion dictionary to determine the emotion recognition result. For example, in an intelligent customer service scenario, if the target user's current input is a consultation question, and this question uses many negative emotion words, it may imply that the user has higher expectations or potential dissatisfaction with the current service; this is a manifestation of implicit needs.
[0065] Optionally, the computer device inputs historical sample sets, emotion recognition results, current input content, the target scenario involved in the current input content, and the corresponding demand tag system into the demand recognition model to obtain the target user's demand description information. In other words, the demand recognition model can learn from historical demand recognition experience based on historical sample sets, and, combined with emotion recognition results, current input content, and the target scenario, determine the most matching target demand tags from the target scenario's demand tag system. Based on these target demand tags, it determines the corresponding quantitative indicators, thereby outputting demand description information.
[0066] In some scenarios, for ease of management, the requirement tags in the requirement tagging system can be encoded, that is, converted into a computer-processable form using appropriate encoding methods. Natural language descriptions combined with numerical encoding can be used to ensure the accuracy and identifiability of the tags. Simultaneously, a hierarchical structure and relationships between tags are established to more comprehensively represent the complexity of implicit requirements. This allows the requirement identification model to better retrieve matching target requirement tags from the requirement tagging system. In some embodiments, the requirement tagging system can also be updated periodically, such as updating and expanding it as the application scenarios of the intelligent agent change and user needs evolve. By collecting new user data, analyzing market trends, and user feedback, new implicit requirement categories can be promptly identified and added to the requirement tagging system, ensuring the timeliness and completeness of the tagging system.
[0067] In this embodiment, the appropriate demand tags in the demand tag system corresponding to the target scenario are evaluated from different aspects by combining historical sample sets, emotion recognition results, current input content, and target scenario, so as to accurately reflect the demand description information of the target user.
[0068] In some embodiments, the demand identification model is a large language model. Based on historical sample sets, sentiment recognition results, current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario, the demand identification model determines the demand description information of the target user, including: determining demand identification prompt text based on historical sample sets, sentiment recognition results, current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario; performing semantic understanding on the demand identification prompt text based on the large language model to obtain the score of each demand tag in the demand tag system; selecting the demand tag corresponding to the highest score as the target demand tag for the target user, and determining the quantitative indicator used to quantify the target demand tag based on the mapping relationship between the demand tag and the quantitative indicator used to quantify the target demand tag; and determining the demand description information of the target user based on the target demand tag and the quantitative indicator used to quantify the target demand tag.
[0069] For example, a computer device acquires a template of a demand identification prompt text, and fills the template with a historical sample set, sentiment recognition results, current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario, thus obtaining the demand identification prompt text. Since each historical sample in the historical sample set also includes a corresponding sentiment recognition result, scenario, and corresponding demand tag, a large language model is used to understand the historical sample set in the demand identification prompt text. Combining the sentiment recognition result, current input content, and target scenario, a score is determined for each demand tag in the demand tag system. For example, the large language model understands the historical sample set in the demand identification prompt text to determine a first score for each demand tag based on the sentiment recognition result, a second score for each demand tag based on keywords related to the demand in the target scenario and current input content, and a final score for each demand tag is determined based on the weights of the first and second scores.
[0070] For example, in order to ensure the accuracy of demand label recognition, it is also necessary to obtain scene information of the intelligent agent's scene in the template, such as by using multi-source data such as sensor data, user input text, and system logs.
[0071] For example, the computer device determines the highest score, and if the highest score is greater than or equal to a score threshold, the demand tag corresponding to the highest score is used as the target demand tag for the target user. Based on the mapping relationship between the demand tag and the quantitative indicator, the quantitative indicator used to quantify the target demand tag is determined. Based on the target demand tag and the quantitative indicator used to quantify the target demand tag, the demand description information of the target user is determined.
[0072] For example, such as Figure 3 The diagram shown is a schematic representation of the steps for determining the requirement description information in one embodiment. Figure 3 The quantitative model in this context is used to derive quantitative indicators after determining demand labels. Since the quantitative indicators are determined based on the mapping relationship between demand labels and quantitative indicators, therefore... Figure 3The quantization model in this context is the large language model. Therefore, after acquiring the current input content, the computer device also acquires the target scene involved in the current input content through the demand capture and quantization unit of the computer device, detects the agent scene information of the agent scene in which the agent is located, identifies the demand-related keywords in the current input content, and acquires the demand tag system corresponding to the target scene. The demand capture and quantization unit uses the emotion recognition model to perform emotion recognition on the current input content, obtains the emotion recognition result, and acquires the historical sample set. The demand capture and quantification unit fills the template with historical sample sets, agent scene information, emotion recognition results, current input content, target scene involved in the current input content, and demand tag system corresponding to the target scene to obtain demand recognition prompt text. It then performs semantic understanding on the demand recognition prompt text using a large language model to obtain the score for each demand tag in the demand tag system. The demand capture and quantification unit selects the demand tag corresponding to the highest score as the target demand tag for the target user and determines the quantification index used to quantify the target demand tag based on the mapping relationship between the demand tag and the quantification index. Finally, based on the target demand tag and the quantification index used to quantify the target demand tag, the demand capture and quantification unit determines the demand description information of the target user and outputs this demand description information.
[0073] In this embodiment, leveraging the understanding capabilities of a large language model, based on historical sample sets, sentiment recognition results, current input content, and the target scenario involved in the current input content, each demand tag in the demand tag system corresponding to the target scenario is scored. The demand tag corresponding to the highest score truly reflects the implicit demand of the current input content, i.e., the target demand tag. At the same time, the corresponding quantitative index can be accurately obtained based on the mapping relationship between demand tags and quantitative indicators. This yields accurate demand description information, ensuring that subsequent processing results can truthfully reflect the agent's processing capabilities, thereby accurately evaluating the agent and ensuring the accuracy and effectiveness of the evaluation.
[0074] In some embodiments, updating the current input content based on scene variables to obtain updated input content includes: creating a variable addition scheme for scene variables; adding scene variables to the current input content according to the variable addition scheme to obtain updated input content.
[0075] The variable addition scheme can be understood as a method for injecting scene variables into the target scene of the current input content. For example, the variable addition scheme indicates the timing and method of scene variable injection to ensure the randomness and rationality of variable injection. Some scene variables can be randomly determined and injected before the target scene begins, or scene variables can be dynamically injected according to certain probabilities and conditions during the target scene to simulate unexpected situations in real-world scenarios. For instance, in an intelligent driving scenario, the scene variable is rain, and the variable addition scheme indicates the timing of the rain and the method of adding it to the target scene.
[0076] For example, after determining the scene variable, the computer device randomly selects one candidate injection method and one candidate method from multiple candidate injection opportunities and methods for the scene variable to create a corresponding variable addition scheme. The computer device then updates the current input content according to this variable addition scheme to obtain the updated input content. For instance, the current input content is a 5-second video of the current traffic conditions. In sunny weather, the traffic light is green. The scene variable is rain, and the corresponding variable addition scheme is that it starts raining lightly at the 5th second of the video. The updated input content is then updated to show that the traffic conditions for the first 4 seconds are sunny with a green light, and the traffic conditions for the 5th second are light rain with a green light.
[0077] For example, such as Figure 4The diagram illustrates the input content update process in one embodiment. After the demand capture and quantification unit of the computer device determines the demand description information, it directly initiates a scene variable construction request to the scene evolution engine of the computer device. The scene evolution engine calls the scene library and random variable generation algorithm to randomly generate scene variables for the target scene based on the current input content and demand description information, and creates a variable addition scheme for the scene variables. According to the variable addition scheme, the scene variables are added to the current input content to obtain the updated input content. After the scene evolution engine sends the updated input content to the agent, it executes step 206 to obtain the evaluation result. To ensure the effectiveness and comprehensiveness of the evaluation, the agent undergoes multiple rounds of evaluation. Therefore, after obtaining the evaluation results, the evaluation round is updated. If the number of evaluation rounds does not reach the preset number, the next round of evaluation is conducted. Specifically, the scene evolution engine, based on the scene variables and evaluation results from the previous round, uses a scene complexity adjustment algorithm to adjust the scene variables from the previous round, resulting in new scene variables. These new scene variables are then used to process the current input content (i.e., the initial input content entered by the target user in the first round of evaluation). The processing result is used to conduct the next round of evaluation until the preset number of rounds is reached. If the evaluation results for all preset rounds indicate that the agent's processing ability is excellent, and the number of rounds indicating reasonable evaluation results is greater than or equal to a preset threshold, then the processing ability is considered acceptable; otherwise, it is considered unacceptable. If the evaluation result of the previous round indicates failure, the scene complexity adjustment algorithm reduces the complexity of the scene variables from the previous round. For example, if it was heavy rain in the previous round, it is adjusted to light rain. If the evaluation results of the previous round indicate that the scenario complexity adjustment algorithm is passed, the complexity of the scenario variables in the previous round will be increased. For example, if it was heavy rain in the previous round, it will be adjusted to torrential rain.
[0078] The scenario library, derived from domain knowledge and historical data, encompasses scenarios across multiple fields, including hotspot prediction, risk management, script recommendation, business negotiation, medical diagnosis, and educational tutoring. For each scenario, the initial state, participating roles and their initial attributes, and the scenario's objectives and rules are defined in detail. The scenario complexity adjustment algorithm adjusts the complexity of scenario variables according to a complexity adjustment strategy. This strategy dynamically adjusts scenario complexity based on the agent's real-time performance. For example, if the agent's evaluation results from previous rounds are reasonable, it indicates excellent performance in the current scenario (high task completion, reasonable decision-making, and rapid response). In this case, the number of scenario variables can be increased, the magnitude of variable changes increased, or more complex task requirements introduced to enhance scenario complexity. Conversely, if the agent performs poorly, indicating unreasonable evaluation results, the scenario complexity can be appropriately reduced to prevent the agent from being unable to cope effectively due to excessive difficulty, thus affecting the accuracy of the evaluation. This scenario complexity adjustment algorithm can utilize machine learning algorithms, such as the Q-Learning algorithm in reinforcement learning, to allow the system to automatically optimize the scenario complexity adjustment strategy through continuous trial and learning, so as to better adapt to the ability levels of different intelligent agents.
[0079] In this embodiment, a variable addition scheme for scene variables is created to inject scene variables into the current input content, thereby simulating the occurrence of sudden events in the real world, taking into account the emergency response of the intelligent agent when sudden events occur, and accurately evaluating the processing capability of the intelligent agent.
[0080] In some embodiments, the agent is evaluated based on the processing results to obtain evaluation results, including: based on the agent's processing results and requirement description information, calling the evaluation model to determine the agent's perception score in the perception dimension, reasoning score in the reasoning dimension, and execution score in the execution dimension; and fusing the perception score, reasoning score, and execution score to obtain an evaluation result that reflects the agent's processing capabilities.
[0081] The perception dimension is used to evaluate the agent's speed and accuracy in capturing needs. In intelligent education scenarios, when a student poses a question (the current input), a scenario variable is added to the question, and the updated input is obtained after post-processing, resulting in an ambiguous question. The evaluation assesses how quickly the agent can capture the student's potential knowledge gaps—the implicit need—and the accuracy of this recognition.
[0082] The reasoning dimension is used to evaluate the agent's reasoning ability, that is, to assess whether the agent accurately responds to the input (updating the input) and obtains a reasonable response. Specifically, after the agent receives changes in scenario variables or identifies implicit needs, it deeply analyzes the agent's adjustment logic to the original strategy (the processing result of the current input) (updating the processing result of the input). By following the agent's decision-making process, including its knowledge invocation, reasoning steps, and decision-making basis, the rationality and effectiveness of its adjusted strategy are evaluated. For example, in an intelligent investment scenario, when the scenario variable of market fluctuation changes, observe how the agent adjusts its portfolio decisions based on market data and its own investment strategy model. The execution dimension is used to evaluate the degree of matching between the needs identified by the agent and the reasoning results determined by the agent. For example, in an e-commerce recommendation scenario, after the agent identifies a user's implicit need for high-value products, it evaluates whether the recommended products truly meet this need in terms of price, performance, etc. In an intelligent healthcare scenario, it examines whether the treatment suggestions given by the doctor-assisted agent based on changes in the patient's condition (scenario evolution) and the patient's concerns about treatment effectiveness (implicit needs) are reasonable and feasible.
[0083] For example, the computer device inputs the agent's processing results and requirement description information into the evaluation model, obtaining the agent's perception score in the perception dimension, reasoning score in the reasoning dimension, and execution score in the execution dimension. The computer device can directly sum the perception score, reasoning score, and execution score to obtain a fusion score (i.e., the final score mentioned above), or it can set different weights for each dimension according to actual needs, and weight the perception score, reasoning score, and execution score according to their respective weights to obtain the fusion score. If the fusion score is lower than a preset score threshold, the evaluation result indicates that the agent failed the evaluation, the agent's processing result is unreasonable, and adjustments to the agent or the complexity of the scenario variables are needed, taking into account external feedback. If the fusion score is higher than or equal to the preset score threshold, the evaluation result indicates that the agent passed the evaluation, i.e., the agent's processing result is reasonable. If the preset number of evaluation rounds is not reached, a new round of evaluation is conducted. The scenario evolution engine adjusts the complexity of the scenario variables from the previous round using a scenario complexity adjustment algorithm, resulting in new scenario variables. These new variables are then used to process the current input (the initial input from the target user in the first round of evaluation). The results of this processing are used to conduct the next round of evaluation until the preset number of rounds is reached. If the evaluation results for all preset rounds indicate excellent processing capability, and the number of rounds indicating reasonable processing capability exceeds a preset threshold, then the processing capability is considered acceptable; otherwise, it is considered unacceptable. The preset number of rounds is less than the preset number of rounds.
[0084] In this embodiment, the processing capability of the agent is comprehensively evaluated from the perspectives of perception, reasoning, and execution, based on the agent's processing results and demand description information, in order to ensure the accuracy of the evaluation.
[0085] In some embodiments, the evaluation model includes a first evaluation model for evaluating the perception dimension, a second evaluation model for evaluating the reasoning dimension, and a third evaluation model for evaluating the execution dimension. Based on the agent's processing results and requirement description information, the evaluation models are invoked to determine the agent's perception score in the perception dimension, reasoning score in the reasoning dimension, and execution score in the execution dimension, respectively. This includes: obtaining requirement identification results and reasoning results from the processing results; determining the perception score based on the requirement identification results and target requirement tags in the requirement description information using the first evaluation model; determining the reasoning score based on the reasoning results and quantitative indicators in the requirement description information using the second evaluation model; and determining the execution score based on the requirement identification results and reasoning results using the third evaluation model.
[0086] After receiving the updated input, the agent performs demand identification and reasoning processes, resulting in demand identification and reasoning results, respectively. The reasoning result refers to the agent's response to the updated input.
[0087] For example, a computer device, based on the demand identification result and the target demand label in the demand description information, obtains a perception score through a first evaluation model. The demand identification result includes the time of the identified demand and the identified demand label. The perception score is determined by the first evaluation model based on the demand identification time (which determines the demand identification speed) and the degree of matching between the demand label identified by the agent and the target demand label in the demand description information. This degree of matching reflects the accuracy of the agent's demand identification. Of course, the demand identification result can also include the scene identification time. Therefore, the first evaluation model can also combine the scene variable identification time to determine the scene variable identification speed. The perception score is obtained by combining the scene variable identification speed, the demand identification speed, and the degree of matching. The faster the identification speed (scene identification speed and demand identification speed) and the higher the degree of matching, the higher the corresponding perception score. The first evaluation model can be built based on a neural network model or a large language model.
[0088] For example, the reasoning result includes the time it takes for the agent to arrive at the reasoning result. The second evaluation model is a large language model. The computer device's reasoning result and the quantitative indicators in the requirement description information determine the reasoning evaluation prompt text. The large language model is then invoked to perform semantic understanding on the reasoning evaluation prompt text to obtain a reasoning score. The reasoning evaluation prompt text indicates that the large language model obtains a reasoning score based on reasoning accuracy, reasoning efficiency, and reasoning innovation. Reasoning accuracy refers to the degree of matching between the reasoning result and the quantitative indicators; the greater the matching degree, the more accurate the reasoning result. For example, updating the input content to show that the road conditions 4 seconds ago are sunny and the green light is on, and the road conditions 5 seconds ago are light rain and the green light is on; the reasoning result is safe passage, and the vehicle speed 4 seconds ago is greater than the vehicle speed 5 seconds ago. The corresponding quantitative indicator is: the goal is safe passage, and the corresponding rules are stop at red lights, go at green lights, and yellow lights, etc. It can be understood that the reasoning result matches the quantitative indicators, meaning the reasoning result is the goal given under the premise of satisfying the rules. Innovative reasoning refers to whether one can flexibly reason based on the scenario variables that update the input content, that is, to provide a novel and effective solution. For example, in the reasoning result: the speed of the car in the first 4 seconds is greater than the speed of the car in the 5th second, which provides a novel and effective solution.
[0089] For example, based on the demand labels and inference results identified in the demand identification results, the computer device uses a third evaluation model to evaluate whether the identified demand labels match the inference results, and obtains an execution score. The better the match, the higher the corresponding execution score. The third evaluation model is built based on a neural network model, or it can be a large language model.
[0090] For example, such as Figure 5 The diagram illustrates the evaluation process in one embodiment. After determining the processing result, the computer device sends the processing result and requirement description information to the evaluation unit of the computer device. The unit then returns to the previous step, which retrieves the requirement identification result, requirement identification result, and reasoning result from the processing result, and continues execution to obtain the evaluation result. Figure 5 The capture speed includes at least the requirement recognition speed mentioned above, and the accuracy is the matching degree mentioned above. Figure 5 The reasoning and demand matching involves determining whether the identified demand labels match the reasoning results. After obtaining the evaluation results, the method also includes introducing new dimensions based on the evaluation results, such as the depth of understanding of the demands (not just identification, but also understanding the reasons behind the demands), and the comprehensive perception ability of multiple variables and implicit demands in complex scenarios, making the evaluation more comprehensive and accurate.
[0091] In this embodiment, based on the demand identification results and target demand labels, the agent's perception ability is evaluated using a first evaluation model, resulting in a score reflecting this ability. Based on the reasoning results and quantitative indicators of the demand description information, the agent's reasoning ability is evaluated using a second evaluation model, resulting in a score reflecting this ability. Based on the demand identification results and reasoning results, the agent's execution ability is evaluated using a second evaluation model, resulting in a score reflecting this ability. Therefore, the agent's processing capability is comprehensively evaluated based on these three dimensions to ensure the accuracy of the evaluation.
[0092] In some embodiments, after step 206 is executed, the evaluation round number is updated (after the first round of evaluation is completed, the evaluation round number is updated to 1; after subsequent rounds of evaluation, it is incremented by 1 based on the previous round's number). If the evaluation round number has not reached the preset number of rounds, the next round of evaluation is performed. That is, the scene evolution engine adjusts the scene variables of the previous round based on the scene variables and evaluation results obtained in the previous round using a scene complexity adjustment algorithm to obtain new scene variables. Based on the new scene variables, the current input content (i.e., the input content first entered by the target user in the first round of evaluation) is processed, and the next round of evaluation is performed based on the processing results until the preset number of rounds is reached. If the evaluation results of the preset number of rounds all indicate that the agent's processing ability is excellent, and if the number of rounds that the evaluation results indicate are reasonable is greater than a preset threshold, then the processing ability is qualified; if it is less than the preset threshold, then the processing ability is unqualified. The preset threshold is less than the preset number of rounds.
[0093] After determining the agent's processing capabilities, a feedback mechanism can be introduced to optimize the entire evaluation process, such as... Figure 6The diagram illustrates the feedback process in one embodiment. The computer device can acquire feedback information through multiple channels, such as in-app feedback channels, online customer service systems, and social media platform monitoring tools, ensuring users can conveniently and quickly submit feedback while covering feedback from different types of users. Specifically, a dedicated feedback entry point is set up within the application to encourage users to submit problems and suggestions during use; social media monitoring tools collect user evaluations and discussions about the agent on social platforms. The computer device sends the feedback information to its feedback collection and processing unit for feedback classification and filtering. This involves classifying and filtering the large amount of collected feedback into different categories such as functional requirements, performance issues, interface design, and requirement-related feedback. A combination of natural language processing technology and manual review is used to filter feedback information related to requirement capture and scenario adaptability, providing targeted data support for subsequent optimization. Then, feedback is quantified and standardized, transforming unstructured user feedback into quantifiable and standardized target feedback information. For user feedback regarding the agent's response to requirements, it can be converted into corresponding satisfaction scores or problem severity scores based on the content and sentiment of the feedback for more intuitive analysis and comparison. Next, the evaluation benchmark (evaluation dimension) is updated based on the target feedback information. That is, a data-driven evaluation benchmark update strategy is formulated based on the target feedback information. If a large number of users report that the intelligent agent has a low accuracy rate in identifying a certain implicit need in a specific scenario, then the evaluation benchmark will place greater emphasis on the identification of that scenario and need, adjust the weight of relevant evaluation indicators (corresponding evaluation dimensions), or add new evaluation indicators (new evaluation dimensions).
[0094] In a specific embodiment, such as Figure 7 The diagram shown illustrates the agent evaluation process in one embodiment. Upon first acquiring the target user's current input, the computer device initiates multiple rounds of evaluation:
[0095] For the first round: Step 7.1: The computer equipment requirements capture and quantification unit performs the following step 7.1.1 to determine the requirements description information corresponding to the captured current input content:
[0096] Step 7.1.1: The demand capture and quantification unit of the computer device acquires a historical sample set, where each historical sample includes historical input content. The agent processes the historical input content to obtain historical processing results. Using an emotion recognition model, the current input content is subjected to emotion recognition to obtain an emotion recognition result. Based on the historical sample set, the emotion recognition result, the current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario, a demand recognition prompt text is determined. The demand recognition prompt text is semantically understood based on the large language model to obtain a score for each demand tag in the demand tag system. The demand tag corresponding to the highest score is selected as the target demand tag for the target user, and a quantification index for quantifying the target demand tag is determined based on the mapping relationship between the demand tag and the quantification index. Based on the target demand tag and the quantification index for quantifying the target demand tag, the demand description information of the target user is determined.
[0097] Step 7.2: The scene evolution engine dynamically generates scene variables to obtain updated input content. See step 7.2.1 below for details:
[0098] Step 7.2.1: Based on the current input content and the requirement description information, randomly generate scene variables for the target scene, and create a variable addition scheme for the scene variables; according to the variable addition scheme, add the scene variables to the current input content to obtain updated input content; the scene involved in the updated input content is the scene obtained by adding the scene variables to the target scene.
[0099] Step 7.3: Send the updated input to the agent so that the agent can process the updated input and obtain the processing result.
[0100] Step 7.4: The evaluation unit of the computer device evaluates the agent from the perspectives of perception, reasoning, and execution to obtain the evaluation results. Specifically, refer to step 7.4.1 below:
[0101] Step 7.4.1: The evaluation unit obtains the requirement identification result and reasoning result from the processing result; based on the requirement identification result and the target requirement tag in the requirement description information, it determines the perception score through the first evaluation model; based on the reasoning result and the quantitative indicator in the requirement description information, it determines the reasoning score through the second evaluation model; based on the requirement identification result and the reasoning result, it determines the execution score through the third evaluation model. The evaluation unit integrates the perception score, reasoning score, and execution score to obtain an evaluation result that reflects the processing capability of the intelligent agent.
[0102] Then, proceed to step 7.5: After completing the first round of evaluation, update the evaluation round to 1 and start the next round of evaluation. That is, for the current round of evaluation that is not the first round of evaluation, obtain the scene variables from the previous round of evaluation. The scene evolution engine adjusts the complexity of the scene variables from the previous round of evaluation. For example, if the evaluation result of the previous round of evaluation indicates a pass, the complexity is increased; if the evaluation result of the previous round of evaluation indicates a fail, the complexity is decreased. Specifically, the adjustment direction is to adjust the degree of the scene variables used in the previous round of evaluation. For example, if the scene variable is rain, the complexity is increased by adjusting the amount of rainfall. Heavy rain or torrential rain increases the complexity, while light rain or no rain decreases it. Alternatively, for the scene variables generated in the previous round of evaluation, the complexity can be increased by adding other scene variables, or the complexity can be decreased by subtracting a certain scene variable from the scene variables of the previous round of evaluation.
[0103] Based on the adjusted scene variables, update the input content of the previous round to obtain the update input content of the current round, and return to steps 7.3-7.4 above to continue execution until the number of evaluation rounds reaches the preset number of rounds. Count the number of rounds in the preset number of rounds where the evaluation results indicate that the rounds have passed. If the evaluation results of all rounds indicate that the rounds have passed, the processing capability is excellent. If the number of rounds indicating that the rounds have passed is equal to the preset number threshold, the processing capability is qualified. If it is less than the preset number threshold, the processing capability is unqualified.
[0104] Step 7.6: After determining the processing capacity, feedback information is collected through the feedback collection unit of the computer equipment to update the process of the requirement capture and quantification unit capturing requirement description information and the process of the scenario evolution engine obtaining updated input content.
[0105] It should be noted that the above process constructs... Figure 7 The adaptive evaluation framework for intelligent agents, based on dynamic scene evolution and the capture of users' (latent) needs, aims to address the limitations of existing intelligent agent evaluations, such as static evaluation sets, fixed task structures, or single interaction modes. This framework is driven by both dynamic scene evolution and the mining of users' latent needs, generating unexpected scene variables and latent need signals in real time to comprehensively evaluate the agent's performance throughout the entire process, from passive response to proactive adaptation. It provides more realistic, comprehensive, and effective evaluation support for the application of intelligent agents in complex and ever-changing real-world scenarios, driving the evolution of intelligent agents from mechanical execution to human-like proactive service. Specifically, the scene evolution engine is built based on real-world scene characteristics, injecting random variables in real time and dynamically adjusting scene complexity. The need capture and quantification unit extracts latent needs from explicit user input and transforms them into quantifiable indicators. The evaluation unit assesses the agent's adaptive capabilities from three dimensions: perception, reasoning, and execution, forming a closed-loop feedback optimization mechanism that dynamically updates the evaluation benchmark based on user feedback.
[0106] In this embodiment, based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, a demand identification model is used to accurately identify the target user's true needs and determine the target user's demand description information. The demand description information includes target demand tags for the target user and quantitative indicators used to quantify the target demand tags. Based on the current input content and demand description information, scenario variables for the target scenario are randomly generated. That is, by combining the target user's true needs, unexpected emergencies that may occur in the target scenario are randomly generated. To this end, the current input content is updated based on the scenario variables to obtain updated input content. The scenario involved in the updated input content is the scenario obtained by adding scenario variables to the target scenario. Thus, the updated input content is based on the current input content with the addition of randomly generated scenario variables. Based on the updated input content, the intelligent agent processes the data to obtain the processing result. Based on the processing result, the intelligent agent is evaluated to obtain the evaluation result. In this way, the processing capability of the intelligent agent under the dynamic addition of unexpected scenario variables in the target scenario can be comprehensively and accurately evaluated, effectively improving the accuracy of the intelligent agent evaluation.
[0107] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0108] Based on the same inventive concept, this application also provides an intelligent agent evaluation device for implementing the intelligent agent evaluation method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more intelligent agent evaluation device embodiments provided below can be found in the limitations of the intelligent agent evaluation method described above, and will not be repeated here.
[0109] In one exemplary embodiment, such as Figure 8As shown, an evaluation device 800 for an intelligent agent is provided, including: an information acquisition module 802, a content update module 804, and an intelligent agent evaluation module 806, wherein:
[0110] The information acquisition module 802 is used to determine the target user's demand description information based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, through a demand identification model. The demand description information includes target demand tags about the target user and quantitative indicators used to quantify the target demand tags.
[0111] The content update module 804 is used to randomly generate scene variables for the target scene based on the current input content and requirement description information, update the current input content based on the scene variables, and obtain updated input content. The scene involved in the updated input content is the scene obtained by adding scene variables to the target scene.
[0112] The agent evaluation module 806 is used to process the updated input content through the agent, obtain the processing result, evaluate the agent based on the processing result, and obtain the evaluation result.
[0113] In some embodiments, the information acquisition module 802 is used to acquire a historical sample set, where each historical sample includes historical input content; the historical input content is processed by an intelligent agent to obtain historical processing results; an emotion recognition model is used to perform emotion recognition on the current input content to obtain emotion recognition results; and based on the historical sample set, the emotion recognition results, the current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario, the demand description information of the target user is determined through a demand recognition model.
[0114] In some embodiments, the demand identification model is a large language model. The information acquisition module 802 is used to determine the demand identification prompt text based on the historical sample set, sentiment recognition results, current input content, the target scene involved in the current input content, and the demand tag system corresponding to the target scene; perform semantic understanding on the demand identification prompt text based on the large language model to obtain the score of each demand tag in the demand tag system; select the demand tag corresponding to the highest score as the target demand tag for the target user, and determine the quantitative index used to quantify the target demand tag based on the mapping relationship between the demand tag and the quantitative index used to quantify the target demand tag; and determine the demand description information of the target user based on the target demand tag and the quantitative index used to quantify the target demand tag.
[0115] In some embodiments, the content update module 804 is used to create a variable addition scheme for scene variables; according to the variable addition scheme, the scene variables are added to the current input content to obtain the updated input content.
[0116] In some embodiments, the agent evaluation module 806 is used to call an evaluation model based on the agent's processing results and requirement description information, and to determine the agent's perception score in the perception dimension, reasoning score in the reasoning dimension, and execution score in the execution dimension; and to fuse the perception score, reasoning score, and execution score to obtain an evaluation result that reflects the agent's processing capability.
[0117] In some embodiments, the evaluation model includes a first evaluation model for evaluating the perception dimension, a second evaluation model for evaluating the reasoning dimension, and a third evaluation model for evaluating the execution dimension. The agent evaluation module 806 is used to obtain the demand identification result and the reasoning result from the processing result; determine the perception score based on the target demand label in the demand identification result and the demand description information through the first evaluation model; determine the reasoning score based on the reasoning result and the quantitative indicator in the demand description information through the second evaluation model; and determine the execution score based on the demand identification result and the reasoning result through the third evaluation model.
[0118] Each module in the aforementioned evaluation device for intelligent agents can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.
[0119] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for evaluating intelligent agents.
[0120] Those skilled in the art will understand that Figure 9The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0121] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0122] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0123] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0125] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0126] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0127] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for evaluating an intelligent agent, characterized in that, The method includes: Based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, the demand description information of the target user is determined through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags. Based on the current input content and the requirement description information, a scene variable for the target scene is randomly generated. The current input content is then updated based on the scene variable to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variable to the target scene. Based on the updated input content, the intelligent agent processes the data to obtain a processing result. Based on the processing result, the intelligent agent is evaluated to obtain an evaluation result.
2. The method according to claim 1, characterized in that, The requirement tag system based on the target user's current input content and the target scenario involved in the current input content, through a requirement identification model, determines the target user's requirement description information, including: A historical sample set is obtained, wherein each historical sample in the historical sample set includes historical input content, and the historical processing result is obtained by the intelligent agent through processing the historical input content; An emotion recognition model is used to perform emotion recognition on the current input content to obtain the emotion recognition result; Based on the historical sample set, the emotion recognition results, the current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario, the demand description information of the target user is determined through the demand recognition model.
3. The method according to claim 2, characterized in that, The demand identification model is a large language model. Based on the historical sample set, the sentiment recognition result, the current input content, the target scenario involved in the current input content, and the demand tag system corresponding to the target scenario, the demand identification model determines the demand description information of the target user, including: Based on the historical sample set, the emotion recognition result, the current input content, the target scene involved in the current input content, and the demand tag system corresponding to the target scene, the demand recognition prompt text is determined; Based on the large language model, semantic understanding is performed on the requirement identification prompt text to obtain the score of each requirement tag in the requirement tag system; The demand tag corresponding to the highest score is selected as the target demand tag for the target user, and the quantitative index used to quantify the target demand tag is determined based on the mapping relationship between the demand tag and the quantitative index. Based on the target demand tags and the quantitative indicators used to quantify the target demand tags, the demand description information of the target user is determined.
4. The method according to claim 1, characterized in that, The step of updating the current input content based on the scene variables to obtain updated input content includes: Create a variable addition scheme for the aforementioned scenario variables; According to the variable addition scheme, the scene variable is added to the current input content to obtain the updated input content.
5. The method according to claim 1, characterized in that, The process of evaluating the agent based on the processing result to obtain the evaluation result includes: Based on the processing results of the intelligent agent and the requirement description information, the evaluation model is invoked to determine the intelligent agent's perception score in the perception dimension, reasoning score in the reasoning dimension, and execution score in the execution dimension, respectively. By integrating perception scores, reasoning scores, and execution scores, an evaluation result reflecting the processing capabilities of the agent is obtained.
6. The method according to claim 5, characterized in that, The evaluation model includes a first evaluation model for evaluating the perception dimension, a second evaluation model for evaluating the reasoning dimension, and a third evaluation model for evaluating the execution dimension. Based on the agent's processing results and the requirement description information, the evaluation models are invoked to determine the agent's perception score in the perception dimension, its reasoning score in the reasoning dimension, and its execution score in the execution dimension, respectively. This includes: Obtain the demand identification result and reasoning result from the processing results; Based on the demand identification results and the target demand tags in the demand description information, the perception score is determined through the first evaluation model; Based on the reasoning results and the quantitative indicators in the requirement description information, the reasoning score is determined through the second evaluation model; Based on the demand identification results and the reasoning results, the execution score is determined through the third evaluation model.
7. An evaluation device for an intelligent agent, characterized in that, The device includes: The information acquisition module is used to determine the target user's demand description information based on the target user's current input content and the demand tag system corresponding to the target scenario involved in the current input content, through a demand identification model. The demand description information includes target demand tags for the target user and quantitative indicators for quantifying the target demand tags. The content update module is used to randomly generate scene variables for the target scene based on the current input content and the requirement description information, and update the current input content based on the scene variables to obtain updated input content. The scene involved in the updated input content is the scene obtained by adding the scene variables to the target scene. The agent evaluation module is used to process the updated input content through an agent to obtain a processing result, and to evaluate the agent based on the processing result to obtain an evaluation result.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.