Method, system and device for evaluating end-side intelligent agent, and evaluation equipment

By analyzing agent behavior data through a multidimensional security assessment framework and a multimodal large language model, the problem of the inability to measure the security and reliability of agents in existing technologies is solved, and a comprehensive assessment and quantitative report generation of agents during task execution is realized.

CN121658320APending Publication Date: 2026-03-13CHINA ACADEMY OF INFORMATION & COMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In existing technologies, the evaluation of an agent's capabilities mainly focuses on task completion rate, which cannot effectively measure its safety and reliability during task execution.

Method used

A multidimensional security assessment framework is provided. By configuring assessment tasks, the intelligent agent is driven to perform interactive operations, monitor and record behavioral data, analyze the behavioral data using a multimodal large language model, and generate an assessment report, including security performance indicators and visualization results.

Benefits of technology

It enables a comprehensive assessment of the security and reliability of intelligent agents during task execution, generates quantitative assessment reports, improves the accuracy and reliability of assessment results, and can identify security risks of intelligent agents from multiple key dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658320A_ABST
    Figure CN121658320A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a method, system and device for evaluating an end-side agent and evaluation equipment, and the method comprises the steps: configuring an evaluation task; the assessment task comprises a preset multi-dimensional security assessment framework and a to-be-tested task instruction; in the end-side equipment environment, driving a target end-side agent to execute an interaction operation according to the task instruction, and monitoring and recording behavior data of the target end-side agent in the process of executing the interaction operation; analyzing behaviors of the target end-side agent according to the behavior data, and determining a safety performance index of the target end-side agent according to the multi-dimensional safety evaluation framework; and according to the behavior data analysis result and the safety performance index, generating an evaluation report of the target end side agent. Through the preset multi-dimensional security assessment framework, the behaviors of the end-side agent can be comprehensively assessed from a plurality of key dimensions, so that the security and reliability of the agent in the task execution process can be more comprehensively identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, such as a method, system, apparatus, and evaluation device for evaluating edge-side intelligent agents. Background Technology

[0002] Currently, with the rapid development of large-scale language model capabilities, edge-side intelligent agents capable of understanding human instructions, perceiving mobile graphical user interfaces, and autonomously executing Android App (application) operations (such as simulating clicks, swipes, and text input) have become a hot topic in research and application. These intelligent agents are expected to achieve a high degree of automation in areas such as mobile office, accessibility functions, and automated testing, thereby greatly improving productivity.

[0003] In related technologies, a method for evaluating the capabilities of an agent is disclosed, which focuses on the agent's task completion rate. This is achieved by setting a series of benchmark tests, including a series of predefined web page tasks, and measuring the agent's performance by assessing whether the agent can ultimately successfully complete the tasks.

[0004] In the process of implementing the embodiments of this disclosure, at least the following problems were found in the related art: The relevant technologies mainly measure whether an intelligent agent can complete a task, but cannot examine the safety and reliability of the intelligent agent during the task execution process.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0007] This disclosure provides a method, system, apparatus, and evaluation device for evaluating edge-side intelligent agents to assess the security and reliability of the intelligent agents during task execution.

[0008] In some embodiments, a method for evaluating an edge-side intelligent agent includes: configuring an evaluation task; the evaluation task includes a preset multidimensional security evaluation framework and task instructions to be tested; in the edge-side device environment, driving the target edge-side intelligent agent to perform interactive operations according to the task instructions, and monitoring and recording the behavioral data of the target edge-side intelligent agent during the execution of the interactive operations; analyzing the behavior of the target edge-side intelligent agent based on the behavioral data, and determining the security performance indicators of the target edge-side intelligent agent according to the multidimensional security evaluation framework; and generating an evaluation report of the target edge-side intelligent agent based on the behavioral data analysis results and the security performance indicators.

[0009] Optionally, the multidimensional security assessment framework includes at least one or more of the following assessment dimensions: assessing whether the task instruction itself contains harmful instructions; assessing whether the agent actively refuses to execute the instruction or actively displays an alarm message on the interface; and assessing whether the agent attempts to execute the key actions in the harmful instructions.

[0010] Optionally, the target edge agent is driven to perform interactive operations according to task instructions, and during the execution of interactive operations, the behavioral data of the target edge agent is monitored and recorded, including: controlling the target edge agent to perform interactive operations in the task instructions based on an automation framework; interactive operations include clicking, inputting, and swiping; performing cyclic detection on the hash value of the foreground App / Activity or UI hierarchy structure of the edge device; and obtaining behavioral data after the target edge agent performs each step of the operation based on the cyclic detection results; behavioral data includes screenshots and UI hierarchy structure.

[0011] Optionally, the method for evaluating the edge agent further includes: recording the behavior data in the form of a log after obtaining the behavior data; and / or ending the acquisition of behavior data when the task instruction has been completed, or an irrecoverable error occurs, or a timeout threshold is reached.

[0012] Optionally, the behavior of the target edge agent is analyzed based on behavioral data, and the security performance indicators of the target edge agent are determined according to a multi-dimensional security assessment framework. This includes: extracting text from screenshots in the behavioral data using a dual-channel OCR module, and intelligently cleaning the extracted text to remove UI noise and obtain the target text; calling a multimodal large language model and guiding it through system prompts to analyze the target text and original behavioral data according to the assessment dimensions in the multi-dimensional security assessment framework; and determining the security performance indicators of the target edge agent based on the analysis results of the multimodal large language model. The security performance indicators include scores, task completion rate, and risk ratio under each security dimension.

[0013] Optionally, an evaluation report for the target edge intelligent agent is generated based on the behavioral data analysis results and security performance indicators, including: generating a visualization report based on the behavioral data analysis results and security performance indicators; the visualization report includes a multi-dimensional security radar chart, a list of risk events, and the reasons for the risk event determination.

[0014] Optionally, the method for evaluating edge-side intelligent agents further includes: after configuring the evaluation task, setting up a benchmark test platform based on the evaluation task and a standardized set of security evaluation test cases; in the benchmark test platform, driving multiple target edge-side intelligent agents to perform interactive operations according to task instructions, and monitoring and recording the behavioral data of each target edge-side intelligent agent during the execution of interactive operations; analyzing the behavior of each target edge-side intelligent agent based on the behavioral data of each target edge-side intelligent agent, and determining the security performance indicators of each target edge-side intelligent agent according to a multi-dimensional security evaluation framework; and generating an evaluation report for each target edge-side intelligent agent based on the behavioral data analysis results and security performance indicators of each target edge-side intelligent agent.

[0015] In some embodiments, a system for evaluating edge-side intelligent agents includes: a front-end module configured to configure an evaluation task; the evaluation task includes a preset multidimensional security evaluation framework and task instructions to be tested; a data acquisition module configured to drive the target edge-side intelligent agent to perform interactive operations according to the task instructions in the edge-side device environment, and monitor and record the behavioral data of the target edge-side intelligent agent during the execution of the interactive operations; an analysis and evaluation module configured to analyze the behavior of the target edge-side intelligent agent based on the behavioral data, and determine the security performance indicators of the target edge-side intelligent agent according to the multidimensional security evaluation framework; and a report generation module configured to generate an evaluation report of the target edge-side intelligent agent based on the behavioral data analysis results and the security performance indicators.

[0016] In some embodiments, the apparatus for evaluating edge agents includes a processor and a memory storing program instructions, the processor being configured to perform the method for evaluating edge agents as described above when the program instructions are executed.

[0017] In some embodiments, the evaluation device includes: an evaluation device body having a communication interface for communicating with an end-side device, the end-side device having an end-side intelligent agent; and a system for evaluating an end-side intelligent agent as described above, or an apparatus for evaluating an end-side intelligent agent as described above, installed on the evaluation device body.

[0018] The method, system, apparatus, and evaluation device for evaluating edge-side intelligent agents provided in the embodiments of this disclosure can achieve the following technical effects: This embodiment of the disclosure achieves full automation of the process from task configuration, data collection, behavior analysis to report generation. By configuring the evaluation task, it can automatically complete subsequent tasks such as driving the intelligent agent to perform tasks, monitoring and recording behavioral data, analyzing data, and generating an evaluation report. Real-time monitoring and recording of the intelligent agent's behavioral data during interactive operations provides a rich information foundation for subsequent analysis, making the evaluation results more accurate and reliable. By analyzing the collected behavioral data, the specific performance of the target edge intelligent agent under various security dimensions can be quantitatively calculated, and an evaluation report containing quantitative scores can be generated. Through a preset multi-dimensional security evaluation framework, the behavior of the edge intelligent agent can be comprehensively evaluated from multiple key dimensions, thereby more comprehensively identifying the security and reliability of the intelligent agent during task execution.

[0019] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0020] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein: Figure 1 This is a schematic diagram of a method for evaluating edge-side intelligent agents provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of a system for evaluating edge-side intelligent agents provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of another method for evaluating edge-side intelligent agents provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of another method for evaluating edge-side intelligent agents provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of another method for evaluating edge-side intelligent agents provided in an embodiment of this disclosure; Figure 6 This is a schematic diagram of a device for evaluating edge-side smart agents provided in an embodiment of this disclosure. Detailed Implementation

[0021] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0022] The terms "first," "second," etc., used in the technical solutions described in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0023] Unless otherwise stated, the term "multiple" means two or more.

[0024] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0025] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0026] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0027] Combination Figure 1 As shown, this disclosure provides a method for evaluating edge-side intelligent agents. The execution subject of this method may be a processor, and the method includes: S101, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional security evaluation framework and the task instructions to be tested.

[0028] S102, in the edge device environment, the processor drives the target edge agent to perform interactive operations according to the task instructions, and monitors and records the behavior data of the target edge agent during the execution of the interactive operations.

[0029] S103, the processor analyzes the behavior of the target edge agent based on behavioral data, and determines the security performance indicators of the target edge agent according to the multi-dimensional security evaluation framework.

[0030] S104, the processor generates an evaluation report for the target edge agent based on behavioral data analysis results and security performance indicators.

[0031] This embodiment of the disclosure achieves full automation of the process from task configuration, data collection, behavior analysis to report generation. By configuring the evaluation task, it can automatically complete subsequent tasks such as driving the intelligent agent to perform tasks, monitoring and recording behavioral data, analyzing data, and generating an evaluation report. Real-time monitoring and recording of the intelligent agent's behavioral data during interactive operations provides a rich information foundation for subsequent analysis, making the evaluation results more accurate and reliable. By analyzing the collected behavioral data, the specific performance of the target edge intelligent agent under various security dimensions can be quantitatively calculated, and an evaluation report containing quantitative scores can be generated. Through a preset multi-dimensional security evaluation framework, the behavior of the edge intelligent agent can be comprehensively evaluated from multiple key dimensions, thereby more comprehensively identifying the security and reliability of the intelligent agent during task execution.

[0032] Based on the above-mentioned methods for evaluating edge agents, combined with Figure 2 As shown, this disclosure provides a system 200 for evaluating edge-side intelligent agents, including: a front-end module 201, a data acquisition module 202, an analysis and evaluation module 203, and a report generation module 204. The front-end module 201 is configured to configure an evaluation task; the evaluation task includes a preset multi-dimensional security evaluation framework and task instructions to be tested. The data acquisition module 202 is configured to drive the target edge-side intelligent agent to perform interactive operations according to the task instructions in the edge-side device environment, and monitor and record the behavioral data of the target edge-side intelligent agent during the execution of the interactive operations. The analysis and evaluation module 203 is configured to analyze the behavior of the target edge-side intelligent agent based on the behavioral data, and determine the security performance indicators of the target edge-side intelligent agent according to the multi-dimensional security evaluation framework. The report generation module 204 is configured to generate an evaluation report of the target edge-side intelligent agent based on the behavioral data analysis results and the security performance indicators.

[0033] In this embodiment, the front-end module 201 can receive user input to configure the evaluation task. The data acquisition module 202 is coupled with the edge device environment and can drive the target edge agent to perform App interaction according to the task configuration. It also integrates a monitoring and recording unit 205 to record its behavior data in real time. The analysis and evaluation module 203 can receive the behavior data, first preprocess the behavior data through the preprocessing unit 206, and then call the multimodal large language model through the large model evaluation unit 207 to analyze the behavior data and calculate various security indicators. Finally, the behavior data analysis results and security performance indicators are stored in the cache management unit 208. The report generation module 204 can integrate and visualize the evaluation results of the analysis and evaluation module 203 to generate the final evaluation report. The system 200 for evaluating edge agents provided in this embodiment realizes a complete closed loop from front-end task configuration, automatic data acquisition on the device, automatic evaluation on the back end, to front-end visualization report through a front-end and back-end separation architecture, thereby achieving automation and standardization of the evaluation process.

[0034] In one specific embodiment, the front-end module 201 provides an interactive interface for the user, implemented using front-end technologies such as React. The core functions of the front-end module 201 include task configuration, task monitoring, and result display. The task configuration function receives user-inputted evaluation task information, such as "test instructions," "target application," and "execution parameters" (e.g., timeout). The task monitoring function displays real-time task progress, operation logs, screenshot previews, OCR-recognized text, and risk warnings. The result display function, upon task completion, calls a back-end interface to retrieve data from the report generation module 204 and uses libraries such as ECharts to visualize the evaluation results as radar charts and lists. The front-end module 201 sends the evaluation task to the back-end data acquisition module 202 via an HTTP POST request (using Axios).

[0035] In one specific embodiment, the data acquisition module 202 is implemented based on Python and the Android automation framework and Android Debug Bridge (ADB). After receiving the evaluation task from the front-end module 201, the data acquisition module 202 is responsible for connecting to the real end-side device (such as an Android device) and executing the core control logic (such as auto_control_hyperos3.py) to drive the target end-side agent to start executing the task. The specific execution actions include: initialization operations (such as device.press("home"), clearing background processes), starting the Agent App (device.app_start), setting task instructions (input_box_.set_text(task)), and clicking send (safe_click(send_btn, ...)).

[0036] In a specific embodiment, the data acquisition module 202 also includes a monitoring and recording unit 205 (e.g., data_collection.py). During the agent's task execution, the monitoring and recording unit 205 starts a main monitoring loop (main_monitor_loop) to continuously detect whether the foreground App / Activity (device.app_current()) changes, or whether the hash value of the UI hierarchy (md5_of_text(xml)) changes, in order to determine when to capture a screenshot (safe_screenshot) and the UI hierarchy (safe_dump), and records these events in the form of a log (append_log). The monitoring and recording unit 205 can also continuously check whether the task is completed (e.g., check_task_done_hyperos3()) or whether the timeout threshold (max_loop_time) has been reached, in order to determine whether to stop acquiring data. All captured behavioral data is associated with a unique task ID and persistently stored.

[0037] In one specific embodiment, the analysis and evaluation module 203 is implemented on the backend (e.g., using the Flask framework based on Python), responsible for exposing API interfaces (e.g., / evaluate, / ocr, / cache) and performing analysis. The analysis and evaluation module 203 is triggered after the data acquisition module 202 completes its test execution. The analysis and evaluation module 203 includes three core units: a preprocessing unit 206, a large model evaluation unit 207, and a cache management unit 208. The preprocessing unit 206, implemented based on preprocessor.py, is responsible for calling dual-channel OCR (PaddleOCR and Pytesseract) to extract text from screenshots (extract_text_from_image) and performing intelligent text cleaning (clean_text) to remove UI noise. The large model evaluation unit 207, implemented based on evaluator.py, is responsible for calling a multimodal large language model (e.g., GPT-4o) and using a detailed system prompt to guide the evaluation. The system prompts that the multimodal large language model must make independent and objective judgments on the three security dimensions ("harmful instruction detection", "agent alarm detection", and "harmful action execution detection"), and must return the results in JSON format. The cache management unit 208 is implemented based on context_manager.py, using the ContextManager class to aggregate the results of the first two units, and persists the evaluation context (input, OCR text, model output, etc.) to a JSON file (save_cache) for frontend requests and traceability.

[0038] In one specific embodiment, the report generation module 204 is responsible for visualizing the quantitative scores and violation details output by the analysis and evaluation module 203. The report generation module 204 is typically integrated into the front-end module 201, via an API (such as / task / result or / cache / ).<task_id> The system retrieves the JSON results saved by the analysis and evaluation module 203 and renders them into a radar chart and a list of risk events.

[0039] Optionally, the multidimensional security assessment framework includes at least one or more of the following assessment dimensions: assessing whether the task instruction itself contains harmful instructions; assessing whether the agent actively refuses to execute the instruction or actively displays an alarm message on the interface; and assessing whether the agent attempts to execute the key actions in the harmful instructions.

[0040] In this embodiment, the multi-dimensional security assessment framework can at least achieve malicious command detection, agent alarm detection, and malicious action execution detection, realizing a comprehensive security assessment framework. This allows for a complete identification of various security risks that edge agents may face during task execution, solving the problem of existing technologies that only focus on task completion rate and neglect security performance under adversarial commands. By detecting the agent's behavior when receiving malicious commands, it determines whether the agent can correctly identify and refuse to execute the malicious commands, or whether it can proactively issue alarm information, thus achieving adversarial security assessment.

[0041] Optionally, the assessment of whether the task instruction itself contains harmful instructions can be implemented in the following ways: predefine a database containing known harmful instructions and compare the task instruction with the database; or, use natural language processing technology to analyze the task instruction to determine whether it contains harmful content; or, combine the context information of the task instruction to analyze whether the instruction may cause harmful behavior.

[0042] Optionally, the assessment of whether the agent actively refuses to execute the instruction or actively displays an alarm message on the interface can be implemented in the following ways: by real-time monitoring of the agent's behavior data during task execution, including screenshots and UI (User Interface) hierarchy, to capture whether the agent displays an alarm message or refuses to execute the task; or, by analyzing the UI hierarchy to determine whether the agent displays an alarm message; or, by using OCR (Optical Character Recognition) technology to extract the text content from the screenshot and analyzing whether the text contains an alarm message or a prompt to refuse to execute the task.

[0043] Optionally, the assessment of whether the agent attempted to execute key actions in a harmful instruction can be implemented in the following ways: by real-time monitoring of the agent's behavioral data during task execution, including screenshots, UI hierarchy, and operation logs, to capture whether the agent attempted to execute key actions in a harmful instruction; or, by predefining a series of key actions and analyzing the collected behavioral data to identify whether the agent attempted to execute these actions; or, by using OCR technology to extract text content from screenshots and analyzing whether the text contains features of key actions.

[0044] Optionally, the edge device environment includes an Android device environment; the target edge intelligent agent includes an Android edge intelligent agent.

[0045] Optionally, the target edge agent is driven to perform interactive operations according to task instructions, and during the execution of interactive operations, the behavioral data of the target edge agent is monitored and recorded, including: controlling the target edge agent to perform interactive operations in the task instructions based on an automation framework; interactive operations include clicking, inputting, and swiping; performing cyclic detection on the hash value of the foreground App / Activity or UI hierarchy structure of the edge device; and obtaining behavioral data after the target edge agent performs each step of the operation based on the cyclic detection results; behavioral data includes screenshots and UI hierarchy structure.

[0046] Combination Figure 3 As shown, this disclosure provides another method for evaluating edge-side agents, including: S301, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional security evaluation framework and the task instructions to be tested.

[0047] S302, in the edge device environment, the processor controls the target edge agent to execute interactive operations in the task instructions based on the automation framework; the interactive operations include clicking, inputting, and swiping.

[0048] S303, the processor performs cyclic detection of the hash value of the foreground App / Activity or UI hierarchy structure of the terminal device.

[0049] S304, the processor obtains the behavioral data of the target side agent after each step of the operation based on the loop detection results; the behavioral data includes screenshots and UI hierarchy structure.

[0050] S305 The processor analyzes the behavior of the target edge agent based on behavioral data and determines the security performance indicators of the target edge agent according to the multi-dimensional security evaluation framework.

[0051] S306: The processor generates an evaluation report for the target edge agent based on behavioral data analysis results and security performance indicators.

[0052] In this embodiment, by cyclically detecting the hash value of the foreground App / Activity or UI hierarchy structure of the terminal device, behavioral data after each step of the agent's operation can be dynamically acquired, ensuring data integrity and real-time performance. During the cyclic detection, if a change is detected in the hash value of the foreground App / Activity or UI hierarchy structure of the terminal device, it indicates that the interface of the terminal device has changed in response to the interaction operation of the target terminal agent, thus allowing the acquisition of behavioral data, capturing screenshots, or recording the new UI hierarchy structure. Unnecessary data collection is avoided; data is only recorded when the interface changes, thereby improving testing efficiency and resource utilization.

[0053] Optionally, the UI hierarchy refers to the hierarchical structure of the application's user interface, usually represented in XML format, which describes the layout and attributes of various UI elements (such as buttons, text boxes, etc.) on the screen.

[0054] Optionally, automation frameworks include Android automation frameworks such as uiautomator2. Combined with Android DebugBridge, these frameworks can control the device via a programming interface to perform actions such as clicking, inputting, and swiping.

[0055] Optionally, the method for evaluating the edge agent further includes: recording the behavior data in the form of a log after obtaining the behavior data; and / or ending the acquisition of behavior data when the task instruction has been completed, or an irrecoverable error occurs, or a timeout threshold is reached.

[0056] In this embodiment, the agent's behavioral data is fully recorded in the form of logs, ensuring data traceability and integrity. The acquisition of behavioral data automatically terminates upon task completion, occurrence of an unrecoverable error, or reaching a timeout threshold, ensuring the efficiency and stability of the testing process.

[0057] Optionally, the behavior of the target edge agent is analyzed based on behavioral data, and the security performance indicators of the target edge agent are determined according to a multi-dimensional security assessment framework. This includes: extracting text from screenshots in the behavioral data using a dual-channel OCR module, and intelligently cleaning the extracted text to remove UI noise and obtain the target text; calling a multimodal large language model and guiding it through system prompts to analyze the target text and original behavioral data according to the assessment dimensions in the multi-dimensional security assessment framework; and determining the security performance indicators of the target edge agent based on the analysis results of the multimodal large language model. The security performance indicators include scores, task completion rate, and risk ratio under each security dimension.

[0058] Combination Figure 4As shown, this disclosure provides another method for evaluating edge-side agents, including: S401, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional security evaluation framework and the task instructions to be tested.

[0059] S402, in the edge device environment, the processor drives the target edge agent to perform interactive operations according to the task instructions, and monitors and records the behavior data of the target edge agent during the execution of the interactive operations.

[0060] The S403 processor uses a dual-channel OCR module to extract text from screenshots in behavioral data, and performs intelligent cleaning on the extracted text to remove UI noise and obtain the target text.

[0061] S404, the processor calls the multimodal large language model, and guides the multimodal large language model to analyze the target text and raw behavioral data according to the evaluation dimensions in the multidimensional security assessment framework through system prompts.

[0062] S405: The processor determines the security performance indicators of the target edge agent based on the analysis results of the multimodal large language model.

[0063] S406: The processor generates an evaluation report for the target edge agent based on behavioral data analysis results and security performance indicators.

[0064] In this embodiment, a dual-channel OCR module efficiently extracts text from screenshots, and intelligent cleaning technology removes UI noise, resulting in accurate target text. The input to the multimodal large language model includes task instructions, target text, and screenshot image data. By comprehensively analyzing this multimodal data and leveraging the powerful capabilities of the multimodal large language model, the behavior of the intelligent agent is comprehensively evaluated. Using the multimodal large language model as the core of the evaluation allows for simultaneous understanding of image and text data, enabling context-aware comprehensive judgment of the agent's behavior, thus solving the problems of non-standard and inefficient manual evaluation. Based on the evaluation dimensions in the multidimensional security evaluation framework, the agent's behavior is quantitatively analyzed to determine its security performance indicators, such as scores under each security dimension, including whether harmful instructions are accurately detected, whether warnings are issued for harmful instructions, whether harmful actions are refused, as well as task completion rate and risk ratio, providing an objective assessment of the agent's security. The system prompts clearly define the judgment criteria for each dimension in the multidimensional security evaluation framework, guiding the multimodal large language model to analyze and ensuring that the model can make accurate judgments according to the predefined evaluation dimensions.

[0065] Optionally, the task completion rate is used to measure the agent's ability to complete tasks.

[0066] Optionally, the risk ratio is used to quantify the proportion of potential risk in an agent's behavior, helping to assess its safety in practical applications. For example, a high risk ratio may indicate that the agent has a higher risk of violating regulations in certain operations.

[0067] Optionally, the dual-channel OCR module includes PaddleOCR and Pytesseract.

[0068] In this embodiment, PaddleOCR is a deep learning-based OCR tool suitable for handling complex scenes and text recognition in multiple languages. Pytesseract is a Python tool based on Tesseract, suitable for quickly implementing simple text extraction functions, especially effective when processing high-quality images. By integrating PaddleOCR and Pytesseract, and supplementing them with intelligent text cleaning logic, the robustness and accuracy of text recognition in complex mobile GUIs are improved.

[0069] Optionally, multimodal large language models include GPT-4o.

[0070] Optionally, an evaluation report for the target edge intelligent agent is generated based on the behavioral data analysis results and security performance indicators, including: generating a visualization report based on the behavioral data analysis results and security performance indicators; the visualization report includes a multi-dimensional security radar chart, a list of risk events, and the reasons for the risk event determination.

[0071] In this embodiment, a visualized evaluation report is generated based on behavioral data analysis results and security performance indicators, making the evaluation results more intuitive and easy to understand. The multi-dimensional security radar chart can display the performance of the target edge agent across different security dimensions, allowing users to quickly understand the agent's performance in these dimensions. The risk event list records in detail the detected risk events and the reasons for their determination, helping users understand the source and basis of each risk event. Through the multi-dimensional security radar chart and the risk event list, the security performance of the target edge agent can be comprehensively presented, providing a comprehensive evaluation perspective.

[0072] Optionally, the method for evaluating edge-side intelligent agents further includes: after configuring the evaluation task, setting up a benchmark test platform based on the evaluation task and a standardized set of security evaluation test cases; in the benchmark test platform, driving multiple target edge-side intelligent agents to perform interactive operations according to task instructions, and monitoring and recording the behavioral data of each target edge-side intelligent agent during the execution of interactive operations; analyzing the behavior of each target edge-side intelligent agent based on the behavioral data of each target edge-side intelligent agent, and determining the security performance indicators of each target edge-side intelligent agent according to a multi-dimensional security evaluation framework; and generating an evaluation report for each target edge-side intelligent agent based on the behavioral data analysis results and security performance indicators of each target edge-side intelligent agent.

[0073] Combination Figure 5 As shown, this disclosure provides another method for evaluating edge-side agents, including: S501, Processor Configuration Evaluation Task; The evaluation task includes a preset multi-dimensional security evaluation framework and the task instructions to be tested.

[0074] The S502 processor sets up a benchmark platform based on the evaluation task and a standardized set of security evaluation use cases.

[0075] In the benchmark test platform, the S503 processor drives multiple target edge agents to perform interactive operations according to task instructions, and monitors and records the behavior data of each target edge agent during the execution of interactive operations.

[0076] S504: The processor analyzes the behavior of each target-side agent based on the behavior data of each target-side agent, and determines the security performance indicators of each target-side agent according to the multi-dimensional security assessment framework.

[0077] S505: The processor generates an evaluation report for each target edge agent based on the behavioral data analysis results and security performance indicators of each target edge agent.

[0078] In this embodiment, a standardized benchmark testing platform is set up according to the evaluation task, providing a unified testing environment for multiple target edge agents. By using a standardized set of security evaluation test cases (such as AndroidAgentSafetyBench.json), multiple target edge agents are simultaneously driven to perform interactive operations within the benchmark testing platform, supporting parallel testing of multiple agents. The behavioral data of each target edge agent is monitored and recorded separately, ensuring the independence and integrity of the data. Based on the behavioral data of each target edge agent, security performance indicators are analyzed and determined, providing personalized evaluation results. Through the benchmark testing platform, a unified testing process and evaluation standards can be provided for multiple agents, ensuring the comparability and consistency of evaluation results.

[0079] The method and system for evaluating edge-side intelligent agents provided in this disclosure can solve the problem that related technologies only focus on task completion rate while neglecting key credibility dimensions such as security, controllability, and reliability; solve the problem that related evaluations cannot measure the behavior of intelligent agents when receiving harmful instructions (such as whether they detect, alert, or execute); solve the problem that related evaluation methods rely on manual review and lack an automated evaluation mechanism that integrates task configuration, automatic execution, multi-dimensional analysis, and result display; and solve the problem that related technologies have difficulty in automating the correlation analysis and quantitative scoring of screenshots (images) and OCR-recognized text.

[0080] The following describes the method for evaluating edge-side intelligent agents provided by this disclosure using two specific implementation scenarios. In the APP interaction task evaluation scenario, this disclosure can verify the reliability and accuracy of Android intelligent agents when performing routine APP interaction tasks (such as simulating clicks, swipes, text input, content reading, etc.). For example, when performing transfer or order placement tasks, it assesses whether the intelligent agent will skip the user confirmation step or incorrectly operate the amount input box. In the adversarial security testing scenario, this disclosure can evaluate the behavior of intelligent agents when receiving harmful instructions (such as "open the APP to search for 'a prohibited item'" or "open the APP to post discriminatory remarks"). Through a multimodal large language model, it can quantitatively determine whether the intelligent agent has the ability to detect (attack_detected), reject (is_agent_warning_present), or execute (is_execute_successful) the harmful instruction.

[0081] Combination Figure 6 As shown, this disclosure provides an apparatus 600 for evaluating edge-side intelligent agents, including a processor 601 and a memory 602. Optionally, the apparatus may further include a communication interface 603 and a bus 604. The processor 601, communication interface 603, and memory 602 can communicate with each other via the bus 604. The communication interface 603 can be used for information transmission. The processor 601 can call logical instructions in the memory 602 to execute the method for evaluating edge-side intelligent agents described in the above embodiments.

[0082] Furthermore, the logic instructions in the aforementioned memory 602 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0083] The memory 602, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 601 executes functional applications and data processing by running the program instructions / modules stored in the memory 602, that is, it implements the method for evaluating edge-side intelligent agents in the above embodiments.

[0084] The memory 602 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 602 may include high-speed random access memory and may also include non-volatile memory.

[0085] This disclosure provides an evaluation device, including: an evaluation device body, and the aforementioned system or apparatus for evaluating edge-side intelligent agents. The evaluation device body is provided with a communication interface for communication connection with an edge-side device, which contains an edge-side intelligent agent. The system or apparatus for evaluating the edge-side intelligent agent is installed in the evaluation device body. The installation relationship described herein is not limited to placement within the evaluation device, but also includes installation connections with other components of the evaluation device, including but not limited to physical connections, electrical connections, or signal transmission connections. Those skilled in the art will understand that the system or apparatus for evaluating edge-side intelligent agents can be adapted to feasible evaluation device bodies to achieve other feasible embodiments.

[0086] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described method for evaluating edge-side intelligent agents.

[0087] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code.

[0088] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the technical solutions described herein. As used in the technical solutions described herein, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed elements. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0089] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0090] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for evaluating edge-side intelligent agents, characterized in that, include: Configure the evaluation task; the evaluation task includes a preset multi-dimensional security evaluation framework and the task instructions to be tested; In the edge device environment, drive the target edge agent to perform interactive operations according to task instructions, and monitor and record the behavior data of the target edge agent during the execution of interactive operations; The behavior of the target edge agent is analyzed based on behavioral data, and the security performance indicators of the target edge agent are determined according to a multi-dimensional security assessment framework. Based on the behavioral data analysis results and security performance indicators, an evaluation report of the target edge intelligent agent is generated.

2. The method according to claim 1, characterized in that, A multidimensional security assessment framework should include at least one or more of the following assessment dimensions: Assess whether the task instructions themselves contain harmful instructions; Assess whether the intelligent agent actively refuses to execute the instruction or actively displays an alarm message on the interface; Assess whether the agent attempted to execute a key action in a harmful instruction.

3. The method according to claim 1, characterized in that, The target edge agent is driven to perform interactive operations according to task instructions, and during the execution of interactive operations, the behavioral data of the target edge agent is monitored and recorded, including: The system controls the target edge agent to execute interactive operations in task instructions based on an automation framework; the interactive operations include clicking, inputting, and swiping. Perform loop detection on the hash values ​​of the foreground App / Activity or UI hierarchy structure of the terminal device; Based on the results of the loop detection, obtain the behavioral data of the target edge agent after each step of the operation; the behavioral data includes screenshots and UI hierarchy structure.

4. The method according to claim 3, characterized in that, Also includes: After obtaining behavioral data, the behavioral data will be recorded in log form; and / or, The acquisition of behavioral data will end once the task instruction has been completed, an unrecoverable error has occurred, or the timeout threshold has been reached.

5. The method according to claim 1, characterized in that, The behavior of the target edge agent is analyzed based on behavioral data, and the security performance indicators of the target edge agent are determined according to a multi-dimensional security assessment framework, including: The text is extracted from screenshots in behavioral data using a dual-channel OCR module. The extracted text is then intelligently cleaned to remove UI noise and obtain the target text. The system invokes a multimodal large language model and guides it through system prompts to analyze the target text and raw behavioral data according to the evaluation dimensions in the multidimensional security assessment framework. Based on the analysis results of the multimodal large language model, the safety performance indicators of the target edge agent are determined; the safety performance indicators include the score, task completion rate and risk ratio under each safety dimension.

6. The method according to claim 1, characterized in that, Based on behavioral data analysis results and security performance indicators, an evaluation report for the target edge agent is generated, including: A visualization report is generated based on the behavioral data analysis results and security performance indicators; the visualization report includes a multi-dimensional security radar chart, a list of risk events, and the reasons for the risk event determination.

7. The method according to any one of claims 1 to 6, characterized in that, Also includes: After configuring the assessment task, set up the benchmark test platform according to the assessment task and the standardized security assessment test case set; In the benchmark testing platform, multiple target edge agents are driven to perform interactive operations according to task instructions, and the behavior data of each target edge agent is monitored and recorded during the execution of interactive operations. The behavior of each target edge agent is analyzed based on the behavioral data of each target edge agent, and the security performance indicators of each target edge agent are determined according to the multi-dimensional security assessment framework. Based on the behavioral data analysis results and security performance indicators of each target edge agent, an evaluation report is generated for each target edge agent.

8. A system for evaluating edge-side intelligent agents, characterized in that, include: The front-end module is configured to perform assessment tasks; the assessment tasks include a preset multi-dimensional security assessment framework and task instructions to be tested. The data acquisition module is configured to drive the target edge agent to perform interactive operations according to task instructions in the edge device environment, and to monitor and record the behavioral data of the target edge agent during the execution of interactive operations. The analysis and evaluation module is configured to analyze the behavior of the target edge agent based on behavioral data and determine the security performance indicators of the target edge agent according to the multi-dimensional security evaluation framework. The report generation module is configured to generate an evaluation report for the target edge agent based on behavioral data analysis results and security performance indicators.

9. An apparatus for evaluating edge-side intelligent agents, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to, when running the program instructions, execute the method for evaluating edge agents as described in any one of claims 1 to 7.

10. An evaluation device, characterized in that, include: The evaluation device itself is equipped with a communication interface for communication with the end-side device, which contains an end-side intelligent agent. The system for evaluating edge-side agents as described in claim 8, or the apparatus for evaluating edge-side agents as described in claim 9, is installed on the evaluation device body.