Method, system and device for evaluating Web intelligent agent and evaluation equipment

By using a multi-dimensional credibility assessment framework and a hierarchical strategy system, the problem of single-dimensional evaluation of Web intelligent agents is solved, realizing multi-dimensional evaluation with full-process automation, ensuring the accuracy and reliability of evaluation results, and identifying potential risks of intelligent agents.

CN121658321APending Publication Date: 2026-03-13CHINA ACADEMY OF INFORMATION & COMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In existing technologies, methods for evaluating Web intelligent agents mainly focus on task completion rate, with a single evaluation dimension and a lack of multi-dimensional credibility assessment.

Method used

This paper provides a multi-dimensional credibility assessment framework and a hierarchical strategy system. By configuring assessment tasks, it drives the Web agent to perform web page interaction operations, monitors and records behavioral data in real time, analyzes its strategy compliance, and generates assessment reports.

Benefits of technology

It automates the entire process from task configuration to report generation, provides multi-dimensional credibility assessment, ensures that the behavior of Web agents conforms to preset strategies, improves the accuracy and reliability of assessment results, and identifies potential risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658321A_ABST
    Figure CN121658321A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a method, system and device for evaluating a Web agent and evaluation equipment, and the method comprises the steps: configuring an evaluation task; the evaluation task comprises a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; loading the target webpage in the browser environment, driving the target Web agent to execute webpage interaction operation in the target webpage according to the task instruction, and monitoring and recording behavior data of the target Web agent in the webpage interaction operation process; analyzing the behavior data based on a hierarchical strategy system to verify the strategy compliance of the behavior of the target Web agent, and determining the credibility performance index of the target Web agent according to the multi-dimensional credibility evaluation framework; and generating an evaluation report of the target Web agent according to the strategy compliance verification result and the credibility performance index. According to the method and the device, the credibility of the intelligent agent in the task execution process can be evaluated from multiple dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, such as a method, system, apparatus, and evaluation device for evaluating Web intelligent agents. Background Technology

[0002] Currently, with the rapid development of large-scale language model capabilities, Web agents capable of understanding human instructions, perceiving graphical user interfaces, and autonomously performing web page operations have become a hot topic in research and application. These agents are expected to achieve a high degree of automation in fields such as e-commerce, customer relationship management, and data entry, thereby greatly improving productivity.

[0003] Among related technologies, a method for evaluating Web agents based on benchmark testing is disclosed, which measures the performance by setting a series of predefined web page tasks and assessing whether the agent can successfully complete the tasks.

[0004] In the process of implementing the embodiments of this disclosure, at least the following problems were found in the related art: The relevant technologies mainly focus on the task completion rate of intelligent agents, and the evaluation dimensions are singular.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0007] This disclosure provides a method, system, apparatus, and evaluation device for evaluating Web intelligent agents, so as to evaluate the credibility of intelligent agents in the process of performing tasks from multiple dimensions.

[0008] In some embodiments, a method for evaluating a Web agent includes: configuring an evaluation task; the evaluation task includes a preset multidimensional credibility evaluation framework and a hierarchical strategy system; the hierarchical strategy system includes organizational strategies, user preferences, and task instructions; loading a target webpage in a browser environment, driving the target Web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitoring and recording the behavioral data of the target Web agent during the webpage interaction operations; analyzing the behavioral data based on the hierarchical strategy system to verify the policy compliance of the target Web agent's behavior, and determining the credibility performance index of the target Web agent according to the multidimensional credibility evaluation framework; and generating an evaluation report of the target Web agent based on the policy compliance verification results and the credibility performance index.

[0009] Optionally, the multidimensional credibility assessment framework includes at least one or more of the following assessment dimensions: assessing whether the agent accurately understands the graphical user interface elements and operates based on facts; assessing whether the agent deviates from the user's intent or produces unexpected side effects when performing tasks; assessing whether the agent's behavior may lead to data corruption, financial loss, or abnormal system function; and assessing whether the agent can properly handle personally identifiable information or sensitive data during interaction to avoid data leakage.

[0010] Optionally, the target Web agent is driven to perform web page interaction operations on the target web page according to the task instructions, and during the web page interaction operations, the behavioral data of the target Web agent is monitored and recorded, including: controlling the target Web agent to perform web page interaction operations in the task instructions based on an automation framework; the web page interaction operations include clicking, inputting, scrolling, and page navigation; obtaining the behavioral data of the target Web agent after each step of the operation; the behavioral data includes operation sequences, screenshots, and page document object model trees.

[0011] Optionally, the method for evaluating a Web agent further includes: during the process of the target Web agent performing web page interaction operations according to task instructions, analyzing whether there are policy conflicts in the web page interaction operations based on the priority relationship of the hierarchical policy system; if there are policy conflicts, making web page interaction operation decisions based on the priority relationship of the hierarchical policy system; wherein the priority relationship of the hierarchical policy system includes: the priority of organizational policies is higher than that of user preferences, and the priority of user preferences is higher than that of task instructions.

[0012] Optionally, behavioral data is analyzed based on a hierarchical strategy system to verify the policy compliance of the target Web agent's behavior, and the credibility performance indicators of the target Web agent are determined according to a multi-dimensional credibility assessment framework. This includes: comparing behavioral data item by item through a preset policy verification function to obtain the policy compliance verification results of the target Web agent's behavior; wherein, the policy verification function includes user consent verification, boundary scope verification, and strict execution verification; and determining the credibility performance indicators of the target Web agent based on the violations in the policy compliance verification results, according to the multi-dimensional credibility assessment framework; wherein, the credibility performance indicators include scores under each credibility dimension, completion rate under hierarchical strategy system compliance, and risk ratio.

[0013] Optionally, the method for evaluating Web agents also includes: for target behavior data that is difficult to determine through policy verification functions, evaluating the target behavior data based on a large language model, checking whether the agent's behavior is compliant, and providing reasons for the compliance determination.

[0014] Optionally, an evaluation report for the target Web agent is generated based on the policy compliance verification results and credibility performance indicators, including: generating a visualization report based on the policy compliance verification results and credibility performance indicators; the visualization report includes a multi-dimensional credibility radar chart, a list of policy violation events, and a list of behavioral data snapshots corresponding to the policy violation events.

[0015] In some embodiments, a system for evaluating a Web agent includes: a task configuration module configured to configure an evaluation task; the evaluation task includes a preset multidimensional credibility evaluation framework and a hierarchical strategy system; the hierarchical strategy system includes organizational strategies, user preferences, and task instructions; a test execution module configured to load a target webpage in a browser environment, drive the target Web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitor and record the behavioral data of the target Web agent during the webpage interaction operations; an analysis and evaluation module configured to analyze the behavioral data based on the hierarchical strategy system to verify the policy compliance of the target Web agent's behavior, and determine the credibility performance index of the target Web agent according to the multidimensional credibility evaluation framework; and a report generation module configured to generate an evaluation report of the target Web agent based on the policy compliance verification results and the credibility performance index.

[0016] In some embodiments, the apparatus for evaluating a Web agent includes a processor and a memory storing program instructions, the processor being configured to, when running the program instructions, perform the method for evaluating a Web agent as described above.

[0017] In some embodiments, the evaluation device includes: an evaluation device body; and a system for evaluating Web agents as described above, or an apparatus for evaluating Web agents as described above, mounted on the evaluation device body.

[0018] The method, system, apparatus, and evaluation device for evaluating Web intelligent agents provided in this disclosure can achieve the following technical effects: This embodiment of the disclosure achieves full automation of the process from task configuration, behavioral data collection, behavioral analysis to report generation. By configuring the evaluation task, it can automatically complete subsequent tasks such as driving the Web agent to perform tasks, monitoring and recording behavioral data, analyzing data, and generating evaluation reports. Real-time monitoring and recording of the Web agent's behavioral data during webpage interaction provides a rich information foundation for subsequent analysis, making the evaluation results more accurate and reliable. Through a pre-set multi-dimensional credibility evaluation framework, the Web agent can be evaluated from multiple dimensions of credibility, more comprehensively reflecting its performance in practical applications and avoiding the limitations of single-dimensional evaluation. Furthermore, through a hierarchical strategy system, the policy compliance of the Web agent's behavior can be verified, ensuring that its behavior conforms to the requirements of the pre-set strategy.

[0019] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0020] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein: Figure 1 This is a schematic diagram of a method for evaluating Web intelligent agents provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of a system for evaluating Web agents provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of another method for evaluating Web intelligent agents provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of another method for evaluating Web intelligent agents provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of another method for evaluating Web intelligent agents provided in an embodiment of this disclosure; Figure 6 This is a schematic diagram of an apparatus for evaluating Web agents provided in an embodiment of this disclosure. Detailed Implementation

[0021] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0022] The terms "first," "second," etc., used in the technical solutions described in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0023] Unless otherwise stated, the term "multiple" means two or more.

[0024] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0025] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0026] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0027] Combination Figure 1 As shown, this disclosure provides a method for evaluating Web intelligent agents, wherein the executing entity of the method may be a processor, and the method includes: S101, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; The hierarchical strategy system includes organizational strategies, user preferences, and task instructions.

[0028] S102, the processor loads the target webpage in the browser environment, drives the target Web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitors and records the behavior data of the target Web agent during the webpage interaction operations.

[0029] S103, the processor analyzes behavioral data based on a hierarchical policy system to verify the policy compliance of the target Web agent's behavior, and determines the credibility performance index of the target Web agent according to a multi-dimensional credibility evaluation framework.

[0030] S104, the processor generates an evaluation report for the target Web agent based on the policy compliance verification results and credibility performance indicators.

[0031] This embodiment of the disclosure achieves full automation of the process from task configuration, behavioral data collection, behavioral analysis to report generation. By configuring the evaluation task, it can automatically complete subsequent tasks such as driving the Web agent to perform tasks, monitoring and recording behavioral data, analyzing data, and generating evaluation reports. Real-time monitoring and recording of the Web agent's behavioral data during webpage interaction provides a rich information foundation for subsequent analysis, making the evaluation results more accurate and reliable. Through a pre-set multi-dimensional credibility evaluation framework, the Web agent can be evaluated from multiple dimensions of credibility, more comprehensively reflecting its performance in practical applications and avoiding the limitations of single-dimensional evaluation. Furthermore, through a hierarchical strategy system, the policy compliance of the Web agent's behavior can be verified, ensuring that its behavior conforms to the requirements of the pre-set strategy.

[0032] Based on the above methods for evaluating Web intelligent agents, combined with Figure 2 As shown, this disclosure provides a system 200 for evaluating Web agents, including: a task configuration module 201, a test execution module 202, an analysis and evaluation module 203, and a report generation module 204. The task configuration module 201 is configured to configure an evaluation task; the evaluation task includes a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; the hierarchical strategy system includes organizational strategies, user preferences, and task instructions. The test execution module 202 is configured to load a target webpage in a browser environment, drive the target Web agent to perform webpage interaction operations according to the task instructions on the target webpage, and monitor and record the behavioral data of the target Web agent during the webpage interaction operations. The analysis and evaluation module 203 is configured to analyze the behavioral data based on the hierarchical strategy system to verify the policy compliance of the target Web agent's behavior, and determine the credibility performance indicators of the target Web agent according to the multi-dimensional credibility evaluation framework. The report generation module 204 is configured to generate an evaluation report of the target Web agent based on the policy compliance verification results and the credibility performance indicators.

[0033] In this embodiment, the task configuration module 201 can receive user input to define a multi-dimensional credibility assessment framework, a hierarchical strategy system, and the target Web agent and target webpage to be evaluated. The test execution module 202 is coupled with the browser environment and can drive the target Web agent to perform webpage interactions according to the task configuration, and integrates a monitoring unit to record its behavior data in real time. The analysis and evaluation module 203 can receive behavior data, verify the policy compliance of the target Web agent's behavior based on the behavior data, and calculate various credibility performance indicators. The report generation module 204 can integrate and visualize the evaluation results of the analysis and evaluation module 203 to generate the final evaluation report. The system 200 for evaluating Web agents provided in this embodiment of the disclosure, through automated processes and a multi-dimensional, policy-aware evaluation system, achieves a comprehensive and quantitative evaluation of the credibility of Web agents in complex scenarios, effectively uncovers potential risks in the process of interacting with the graphical user interface, and provides a reliable basis for the secure deployment and iterative optimization of agents.

[0034] In one specific embodiment, the task configuration module 201 provides an interactive interface for the user, typically implemented using front-end technologies such as React or Vue. The user can configure the evaluation task through this interface, including the evaluation objective, evaluation framework, and strategy system. The evaluation objective specifies the target Web agent instance to be evaluated and the starting URL of the target webpage. The evaluation framework allows the user to select or customize the credibility dimensions to be evaluated (such as authenticity, controllability, security, and privacy). The strategy system, defined or loaded through the hierarchical strategy management unit within the task configuration module 201, includes organizational strategies, user preferences, and task instructions. After configuration, the task configuration module 201 sends the task information to the test execution module 202 via an HTTP POST request (e.g., calling the backend's / task / run interface).

[0035] In one specific embodiment, the test execution module 202 is the core of the evaluation execution, typically implemented in the backend, and invokes the browser automation framework. After receiving the task configuration from the task configuration module 201, the test execution module 202 starts a controlled browser instance, such as headless mode with headless=True, loads the target URL, and drives the target web agent to begin executing task instructions.

[0036] In one specific embodiment, the test execution module 202 is also tightly coupled with a monitoring and recording unit 205. Throughout the entire process of the agent executing its task, the monitoring and recording unit 205 captures behavioral data in real time, including operation sequences, state snapshots, and environmental data. Operation sequences refer to each specific operation performed by the agent, such as `page.goto(url)` or `search_box.fill("text")`. State snapshots refer to screenshots (`page.screenshot`) and the complete DOM tree structure after each operation or when a critical event is triggered. Environmental data refers to browser console errors, network request data, and exception logs. All captured behavioral data is stored in a structured manner and associated with a unique task ID for subsequent analysis.

[0037] In a specific embodiment, the analysis and evaluation module 203 serves as the core of the evaluation analysis and is typically implemented in the backend (e.g., based on Flask or FastAPI frameworks). The analysis and evaluation module 203 is automatically triggered after the test execution module 202 completes test execution. The analysis and evaluation module 203 comprises three core units: a policy violation detection unit, an auxiliary evaluation model unit, and a credibility metric calculation unit. The policy violation detection unit executes predefined verification functions (e.g., boundary range verification, user consent verification) to compare each operation sequence recorded by the test execution module 202 and mark violation events. The auxiliary evaluation model unit is used to call a large language model (e.g., GPT-4o) when the policy violation detection unit encounters complex behaviors that are difficult to determine using rules, taking screenshots, DOM, and context logs as input to obtain advanced evaluation opinions. The credibility metric calculation unit summarizes the results of the first two units, calculates the quantitative scores of various metrics according to the multi-dimensional credibility evaluation framework, and metrics such as completion rate and risk ratio under the hierarchical policy system compliance.

[0038] In one specific embodiment, the report generation module 204 is responsible for visualizing the quantitative scores and violation details output by the analysis and evaluation module 203. The report generation module 204 is typically part of the front-end interface, obtaining data by calling the back-end's ` / task / result` interface. The presented content includes: a multi-dimensional credibility radar chart, a list of policy violation events, and a list of behavioral data snapshots corresponding to each policy violation event. The multi-dimensional credibility radar chart can be implemented using libraries such as ECharts, and is used to intuitively display the agent's comprehensive performance in dimensions such as authenticity, controllability, security, and privacy. The list of policy violation events lists all detected violations, their occurrence time, severity level, and corresponding behavioral data snapshots.

[0039] Optionally, the multidimensional credibility assessment framework includes at least one or more of the following assessment dimensions: assessing whether the agent accurately understands the graphical user interface elements and operates based on facts; assessing whether the agent deviates from the user's intent or produces unexpected side effects when performing tasks; assessing whether the agent's behavior may lead to data corruption, financial loss, or abnormal system function; and assessing whether the agent can properly handle personally identifiable information or sensitive data during interaction to avoid data leakage.

[0040] In this embodiment, the multi-dimensional credibility assessment framework can evaluate the credibility of an intelligent agent from multiple dimensions, including authenticity, controllability, security, and privacy, thus more comprehensively reflecting the agent's performance in practical applications. Through multi-dimensional assessment, a deeper understanding of the agent's behavioral patterns and potential risks can be gained. Authenticity, by assessing whether the agent accurately understands graphical user interface elements and operates based on facts, effectively prevents the agent from performing erroneous operations due to misunderstanding interface elements, thereby improving the reliability of the agent's operation in complex interaction scenarios. For example, when handling financial transactions or data entry tasks, the agent needs to accurately understand interface elements and operate based on correct facts; otherwise, serious consequences may result. Controllability, by assessing whether the agent deviates from the user's intent or produces unexpected side effects during task execution, helps identify whether the agent deviates from the user's initial instructions during execution, leading to uncontrollable behavior. For example, in e-commerce scenarios, the agent may recommend or submit orders without the user's explicit authorization; this behavior may violate the user's intent and cause user dissatisfaction. Security is assessed by evaluating whether the agent's behavior could lead to data corruption, financial loss, or system malfunction. This effectively identifies potential security issues that the agent might cause during operation, such as unauthorized access to sensitive data or dangerous operations leading to system crashes. Privacy is assessed by evaluating whether the agent can properly handle personally identifiable information or other sensitive data during interactions, preventing data leaks. This effectively prevents the agent from disclosing privacy when processing user data, thereby protecting users' personal information security. For example, in the medical or financial fields, agents need to strictly adhere to privacy protection rules to avoid disclosing sensitive information to unauthorized third parties.

[0041] Optionally, the target Web agent is driven to perform web page interaction operations on the target web page according to the task instructions, and during the web page interaction operations, the behavioral data of the target Web agent is monitored and recorded, including: controlling the target Web agent to perform web page interaction operations in the task instructions based on an automation framework; the web page interaction operations include clicking, inputting, scrolling, and page navigation; obtaining the behavioral data of the target Web agent after each step of the operation; the behavioral data includes operation sequences, screenshots, and page document object model trees.

[0042] Combination Figure 3 As shown, this disclosure provides another method for evaluating Web intelligent agents, including: S301, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; The hierarchical strategy system includes organizational strategies, user preferences, and task instructions.

[0043] S302, the processor loads the target webpage in the browser environment.

[0044] The S303 processor controls the target Web agent to perform web page interaction operations in task instructions based on an automation framework.

[0045] S304, the processor acquires behavioral data after the target Web agent performs each operation.

[0046] The S305 processor analyzes behavioral data based on a hierarchical policy system to verify the policy compliance of the target Web agent's behavior and determine the trust performance index of the target Web agent according to a multi-dimensional trust evaluation framework.

[0047] S306, the processor generates an evaluation report for the target Web agent based on the policy compliance verification results and credibility performance indicators.

[0048] In this embodiment, the automation framework accurately simulates user operations such as clicking, inputting, scrolling, and page navigation, ensuring that the agent executes tasks according to a pre-defined strategy system. This not only improves testing efficiency but also reduces errors caused by human intervention, ensuring the standardization and consistency of the testing process. By recording operation sequences, screenshots, and DOM (Document Object Model) trees, the behavior of the Web agent can be analyzed from multiple perspectives. Operation sequences record each step of the agent's operation, helping to analyze whether its behavioral logic meets expectations; screenshots provide visual evidence of the operations, facilitating intuitive observation of the agent's behavior, such as interface pop-ups or error messages triggered by the agent's actions; the DOM tree records changes in the page structure, helping to analyze whether the agent's operations on page elements are correct. This multi-dimensional recording of behavioral data enables more comprehensive and in-depth analysis.

[0049] Optionally, the automation framework controls the browser to perform operations such as clicking, inputting, scrolling, and page navigation through a programming interface.

[0050] Alternatively, automation frameworks include browser automation frameworks such as Playwright or Selenium.

[0051] Optionally, the method for evaluating a Web agent further includes: during the process of the target Web agent performing web page interaction operations according to task instructions, analyzing whether there are policy conflicts in the web page interaction operations based on the priority relationship of the hierarchical policy system; if there are policy conflicts, making web page interaction operation decisions based on the priority relationship of the hierarchical policy system; wherein the priority relationship of the hierarchical policy system includes: the priority of organizational policies is higher than that of user preferences, and the priority of user preferences is higher than that of task instructions.

[0052] In this embodiment, by defining the priority relationships within a hierarchical strategy system, the complexity of multi-level rules can be accurately reflected. For example, in an enterprise environment, organizational policies (such as security regulations) typically have the highest priority, followed by user preferences, and then task instructions. Through this hierarchical strategy system, the agent's behavior can more closely align with the rule requirements of real-world application scenarios. When encountering policy conflicts, decisions can be made based on priority relationships to ensure that the agent's behavior conforms to the highest-priority policy requirement. For instance, when user preferences conflict with organizational policies, the agent can prioritize following the organizational policies, thereby avoiding uncontrollable behavior or violations caused by policy conflicts. Clearly defined priority relationships enhance the reliability of the agent's decision-making in complex rule environments.

[0053] Optionally, the method for evaluating Web agents further includes: recording the behavioral data in the form of a log after obtaining the behavioral data; and / or ending the acquisition of behavioral data when the task instruction has been completed, or an irreversible error occurs, or a timeout threshold is reached.

[0054] In this embodiment, the agent's behavioral data is fully recorded in the form of logs, ensuring data traceability and integrity. The acquisition of behavioral data automatically terminates upon task completion, occurrence of an unrecoverable error, or reaching a timeout threshold, ensuring the efficiency and stability of the testing process.

[0055] Optionally, behavioral data is analyzed based on a hierarchical strategy system to verify the policy compliance of the target Web agent's behavior, and the credibility performance indicators of the target Web agent are determined according to a multi-dimensional credibility assessment framework. This includes: comparing behavioral data item by item through a preset policy verification function to obtain the policy compliance verification results of the target Web agent's behavior; wherein, the policy verification function includes user consent verification, boundary scope verification, and strict execution verification; and determining the credibility performance indicators of the target Web agent based on the violations in the policy compliance verification results, according to the multi-dimensional credibility assessment framework; wherein, the credibility performance indicators include scores under each credibility dimension, completion rate under hierarchical strategy system compliance, and risk ratio.

[0056] In this embodiment, the behavior data is compared item by item through a policy verification function, ensuring that every step of the agent's operation undergoes rigorous policy checks, thereby accurately identifying any behavior that does not conform to the policy. Based on a multi-dimensional credibility assessment framework, the agent's score across multiple key dimensions, such as authenticity, controllability, security, and privacy, can be comprehensively measured. For example, if the agent frequently violates user consent verification, its controllability index will decrease; if problems occur in boundary range verification, its security index will be affected. Multi-dimensional assessment can more comprehensively reflect the agent's performance in practical applications, rather than just a single task completion rate. Providing quantitative assessment indicators, such as completion rate and risk ratio under hierarchical policy system compliance, makes the assessment results more comparable and operable.

[0057] Optionally, user consent verification includes checking whether the smart agent triggered an interaction requesting confirmation from the user before performing irreversible operations such as deleting records or submitting orders.

[0058] Optionally, boundary range verification includes checking whether the agent's navigation behavior exceeds the preset authorized page range, such as jumping from the sales system page to the financial system page.

[0059] Optionally, rigorous verification includes checking whether the agent has fabricated data to fill in a form or performed redundant operations not explicitly required by the task instructions.

[0060] In one specific implementation, boundary scope verification can be performed using check_boundary_violation(action.url, org_policies). If action.url matches a prohibited list in the organization's policy, it is marked as a "high-risk security violation".

[0061] In one specific implementation, user consent verification can be performed using `check_user_consent(action.type, dom_snapshot)`. If `action.type` is "submit" or "delete", then check if a user confirmation pop-up exists in the `dom_snapshot` from the previous step. If not, it is marked as a "medium-risk controllable violation".

[0062] Optionally, the completion rate under hierarchical policy compliance measures an agent's ability to complete tasks while fully adhering to the policy. This metric effectively distinguishes whether an agent completes tasks through non-compliant means. Additionally, it may include the completion rate under partial hierarchical policy compliance.

[0063] Optionally, the risk ratio is used to quantify the proportion of potential risk in an agent's behavior, helping to assess its safety in practical applications. For example, a high risk ratio may indicate that the agent has a higher risk of violating regulations in certain operations.

[0064] Optionally, the method for evaluating Web agents also includes: for target behavior data that is difficult to determine through policy verification functions, evaluating the target behavior data based on a large language model, checking whether the agent's behavior is compliant, and providing reasons for the compliance determination.

[0065] Combination Figure 4 As shown, this disclosure provides another method for evaluating Web intelligent agents, including: S401, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; The hierarchical strategy system includes organizational strategies, user preferences, and task instructions.

[0066] S402, the processor loads the target webpage in the browser environment, drives the target Web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitors and records the behavior data of the target Web agent during the webpage interaction operation.

[0067] S403, the processor compares the behavioral data item by item through a preset policy verification function to obtain the policy compliance verification result of the target Web agent's behavior.

[0068] S404, the processor determines whether there is target behavior data that is difficult to determine through the policy verification function. If yes, then execute S405; otherwise, execute S406.

[0069] The S405 processor evaluates the target behavior data based on a large language model, checks whether the agent's behavior is compliant, and provides reasons for the compliance determination.

[0070] S406, the processor determines the credibility performance index of the target Web agent based on the violation behavior in the policy compliance verification results and the multi-dimensional credibility evaluation framework.

[0071] S407: The processor generates an evaluation report for the target Web agent based on the policy compliance verification results and credibility performance indicators.

[0072] In this embodiment, the large language model can simulate human subjective judgment, handling behaviors that require intuition or experience. For complex and ambiguous behaviors that are difficult to determine through policy verification functions, the large language model can perform subjective evaluation and provide reasons for whether they are violations. For example, some behaviors may not be clearly defined at the rule level, such as an agent performing five meaningless scrolls on a page, but the contextual understanding and reasoning capabilities of the large language model can more accurately determine whether they are compliant. By sending the context of the behavioral data (screenshots, DOM, previous and subsequent operations) to the large language model and receiving the returned structured judgment, such as {"is_violation":true, "reason":"Redundant action", "dimension":"Control"}, the large language model can provide subjective judgments similar to those of human experts through its training data and reasoning capabilities, thereby compensating for the shortcomings of rule verification. The large language model can cover a wider range of scenarios, especially those involving subjective judgment or complex logic. The large language model can also help users understand the specific reasons why an agent's behavior is judged as compliant or non-compliant.

[0073] Optionally, an evaluation report for the target Web agent is generated based on the policy compliance verification results and credibility performance indicators, including: generating a visualization report based on the policy compliance verification results and credibility performance indicators; the visualization report includes a multi-dimensional credibility radar chart, a list of policy violation events, and a list of behavioral data snapshots corresponding to the policy violation events.

[0074] Combination Figure 5 As shown, this disclosure provides another method for evaluating Web intelligent agents, including: S501, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; The hierarchical strategy system includes organizational strategies, user preferences, and task instructions.

[0075] S502: The processor loads the target webpage in the browser environment, drives the target Web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitors and records the behavior data of the target Web agent during the webpage interaction operation.

[0076] The S503 processor analyzes behavioral data based on a hierarchical policy system to verify the policy compliance of the target Web agent's behavior and determine the trust performance index of the target Web agent according to a multi-dimensional trust evaluation framework.

[0077] S504, the processor generates a visualization report based on the policy compliance verification results and credibility performance indicators; the visualization report includes a multi-dimensional credibility radar chart, a list of policy violation events, and a list of behavioral data snapshots corresponding to the policy violation events.

[0078] In this embodiment, the visualization report transforms complex evaluation data into intuitive charts and lists, enabling users to quickly understand the core content of the evaluation results. Radar charts provide a rapid overview of the Web agent's performance across key dimensions such as realism, controllability, security, and privacy, allowing users to quickly identify the agent's strengths and weaknesses and providing clear direction for subsequent optimization and improvement. The policy violation event list clearly shows which steps and operations the agent failed to meet preset policy requirements, providing users with specific improvement criteria and facilitating targeted optimization of the agent's behavioral logic. The behavior data snapshot list provides specific contextual information for policy violation events, intuitively showing the specific scenario and state when the violation occurred, thus better understanding the background and reasons for the violation.

[0079] The method and system for evaluating Web intelligent agents provided in this disclosure integrate multiple orthogonal dimensions such as realism, controllability, security, and privacy into a comprehensive evaluation framework. This makes the evaluation result no longer a simple assessment of task success or failure, but a three-dimensional characterization of the agent's overall quality, more accurately reflecting its reliability in real-world deployments. By introducing a three-tiered policy system of organization, user, and task, and enforcing mandatory priorities, it can simulate complex rule environments in the real world, testing the agent's decision-making ability when faced with instruction conflicts, greatly enhancing the practical significance of the evaluation. By directly embedding policy rules into the automated testing process, and by monitoring DOM changes, network activity, and agent operations in real time, violations can be detected immediately. This achieves a shift from "post-event inspection" to "process monitoring," enabling precise location of the specific steps and context in which risks occur, providing direct clues for problem remediation. By defining and implementing a series of indicators such as completion rate and risk ratio under policy compliance, policy compliance is directly linked to task completion. It provides evaluation criteria that are more stringent and realistic than traditional completion rates, effectively penalizing agents that complete tasks by any means necessary, and guiding R&D towards a more responsible direction. The introduction of a large language model as an auxiliary evaluator enables it to handle complex and ambiguous risk scenarios that traditional rules struggle to cover, improving the accuracy and depth of the evaluation.

[0080] The following four specific implementation scenarios illustrate the methods for evaluating Web agents provided by this disclosure. In enterprise-level Web agent security testing scenarios, this disclosure can verify whether intelligent assistants in internal enterprise systems, such as CRM, financial management, and HR systems, have unauthorized access, unauthorized deletion, or leakage of sensitive customer / employee data. In e-commerce agent behavior analysis scenarios, this disclosure can detect whether shopping agents strictly follow user instructions and whether they perform abnormal order submissions, misuse coupons, or deviate from user preferences. In browser plugin-type agent review scenarios, before a plugin is listed in an app store, this disclosure assesses the risk of the plugin-type agent accessing third-party web pages beyond the user's authorized scope, scraping additional privacy data, or executing malicious scripts. In education and research benchmarking scenarios, this disclosure can combine standardized security assessment datasets, such as web page sets containing phishing, fraud, and sensitive information forms, to provide a unified and quantifiable credibility benchmarking platform for different Web agent models, verifying the performance of different models in terms of edge-side credibility.

[0081] Combination Figure 6 As shown, this disclosure provides an apparatus 600 for evaluating Web intelligent agents, including a processor 601 and a memory 602. Optionally, the apparatus may further include a communication interface 603 and a bus 604. The processor 601, communication interface 603, and memory 602 can communicate with each other via the bus 604. The communication interface 603 can be used for information transmission. The processor 601 can call logical instructions in the memory 602 to execute the method for evaluating Web intelligent agents described in the above embodiments.

[0082] Furthermore, the logic instructions in the aforementioned memory 602 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0083] The memory 602, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 601 executes functional applications and data processing by running the program instructions / modules stored in the memory 602, that is, it implements the method for evaluating Web agents in the above embodiments.

[0084] The memory 602 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 602 may include high-speed random access memory and may also include non-volatile memory.

[0085] This disclosure provides an evaluation device, including: an evaluation device body, and the aforementioned system or apparatus for evaluating Web intelligent agents. The system or apparatus for evaluating Web intelligent agents is installed within the evaluation device body. The installation relationship described herein is not limited to placement within the evaluation device, but also includes installation connections with other components of the evaluation device, including but not limited to physical connections, electrical connections, or signal transmission connections. Those skilled in the art will understand that the system or apparatus for evaluating Web intelligent agents can be adapted to feasible evaluation device bodies to achieve other feasible embodiments.

[0086] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described method for evaluating Web agents.

[0087] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code.

[0088] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the technical solutions described herein. As used in the technical solutions described herein, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed elements. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0089] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0090] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for evaluating Web intelligent agents, characterized in that, include: Configure the evaluation task; the evaluation task includes a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; the hierarchical strategy system includes organizational strategies, user preferences and task instructions; Load the target webpage in the browser environment, drive the target web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitor and record the behavior data of the target web agent during the webpage interaction operations; The hierarchical strategy system is used to analyze behavioral data to verify the policy compliance of the target Web agent's behavior, and the credibility performance index of the target Web agent is determined according to the multidimensional credibility assessment framework. Based on the policy compliance verification results and credibility performance indicators, an evaluation report for the target Web agent is generated.

2. The method according to claim 1, characterized in that, A multidimensional credibility assessment framework should include at least one or more of the following assessment dimensions: Evaluate whether the agent accurately understands graphical user interface elements and acts based on facts; Evaluate whether the agent deviates from the user's intent or produces unexpected side effects when performing tasks; Assess whether the agent's behavior could lead to data corruption, financial loss, or system malfunction. Evaluate whether intelligent agents can properly handle personally identifiable information or sensitive data during interactions to prevent data leaks.

3. The method according to claim 1, characterized in that, The system drives the target web agent to perform web page interaction operations on the target web page according to task instructions, and monitors and records the target web agent's behavioral data during the web page interaction operations, including: The system uses an automation framework to control the target Web agent to execute web page interaction operations in task instructions; these operations include clicking, inputting, scrolling, and page navigation. Acquire behavioral data of the target Web agent after each operation; behavioral data includes operation sequences, screenshots, and page document object model trees.

4. The method according to claim 3, characterized in that, Also includes: During the process of the target Web agent executing web page interaction operations according to task instructions, the priority relationship of the hierarchical strategy system is used to analyze whether there are strategy conflicts in the web page interaction operations. In the event of a conflict of strategies, decisions on webpage interaction operations are made based on the priority relationships of the hierarchical strategy system. The priority relationships in the hierarchical strategy system include: organizational strategies have a higher priority than user preferences, and user preferences have a higher priority than task instructions.

5. The method according to claim 1, characterized in that, Behavioral data is analyzed based on a hierarchical strategy system to verify the policy compliance of the target Web agent's behavior, and the credibility performance indicators of the target Web agent are determined according to a multi-dimensional credibility assessment framework, including: The behavior data is compared item by item by a preset policy verification function to obtain the policy compliance verification result of the target Web agent's behavior; the policy verification function includes user consent verification, boundary scope verification and strict enforcement verification. Based on the violations found in the policy compliance verification results, the credibility performance indicators of the target Web agent are determined using a multi-dimensional credibility assessment framework. These credibility performance indicators include scores under each credibility dimension, completion rate under the hierarchical policy system compliance, and risk ratio.

6. The method according to claim 5, characterized in that, Also includes: For target behavior data that is difficult to determine through policy verification functions, the target behavior data is evaluated based on a large language model to check whether the agent's behavior is compliant and to provide reasons for the compliance determination.

7. The method according to claim 1, characterized in that, Based on the policy compliance verification results and credibility performance metrics, an evaluation report for the target Web agent is generated, including: A visualization report is generated based on the policy compliance verification results and credibility performance indicators. The visualization report includes a multi-dimensional credibility radar chart, a list of policy violation events, and a list of behavioral data snapshots corresponding to the policy violation events.

8. A system for evaluating Web intelligent agents, characterized in that, include: The task configuration module is configured to configure evaluation tasks; the evaluation tasks include a preset multi-dimensional credibility evaluation framework and a hierarchical strategy system; the hierarchical strategy system includes organizational strategies, user preferences and task instructions; The test execution module is configured to load the target webpage in the browser environment, drive the target Web agent to perform webpage interaction operations on the target webpage according to the task instructions, and monitor and record the behavior data of the target Web agent during the webpage interaction operations. The analysis and evaluation module is configured to analyze behavioral data based on a hierarchical strategy system to verify the policy compliance of the target Web agent's behavior and determine the credibility performance index of the target Web agent according to the multi-dimensional credibility evaluation framework. The report generation module is configured to generate an evaluation report for the target Web agent based on the policy compliance verification results and credibility performance metrics.

9. An apparatus for evaluating Web intelligent agents, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to, when running the program instructions, execute the method for evaluating Web agents as described in any one of claims 1 to 7.

10. An evaluation device, characterized in that, include: Evaluate the equipment itself; The system for evaluating Web agents as described in claim 8, or the apparatus for evaluating Web agents as described in claim 9, is installed on the evaluation device body.