Intelligent agent network access evaluation method and system based on automatic label generation and multi-dimensional evaluation
By adopting an agent network entry evaluation method based on automated tag generation and multi-dimensional assessment, the problems of standardization, low efficiency and difficult diagnosis in agent evaluation are solved. It realizes the standardization, automation and efficient management of agent evaluation, ensures the quality and reliability of agents, and is suitable for agent management and scheduling in multi-agent collaborative platforms.
Patent Information
- Application Number
- CN202511117886.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for evaluating intelligent agents suffer from fragmentation and lack of standardization, low evaluation efficiency and high cost, black-box characteristics leading to diagnostic difficulties, and insufficient verification of generalization in the real world, making it difficult to achieve efficient and reliable evaluation and management of intelligent agents.
An intelligent agent network access evaluation method based on automated tag generation and multi-dimensional assessment is adopted. It uses a lightweight large language model to automatically generate multi-dimensional tags, collects evaluation indicators through simulated interaction and testing, and generates an evaluation report by combining automated scoring and problem diagnosis, and updates the intelligent agent profile to ensure that the intelligent agent complies with the network access specifications.
It achieves standardization and automation of intelligent agent evaluation, objectivity and multi-dimensionality of evaluation results, ensures the quality and reliability of intelligent agents, simplifies the network access process, improves management efficiency, and supports efficient operation in resource-constrained environments.
Smart Images

Figure CN120950407A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to an intelligent agent evaluation, management and scheduling technology, which is particularly suitable for the standardized and automated evaluation, profile construction and quality management of heterogeneous intelligent agents in a multi-agent collaborative platform. Background Technology
[0002] While AI agent technology is developing rapidly, with various agents based on Large Language Models (LLMs) emerging in large numbers, comprehensive, accurate, and reliable evaluation and effective management still face many challenges. Existing technologies have the following shortcomings in agent network access evaluation and quality management: 1. Fragmentation and Lack of Standardization: There is a lack of unified and authoritative evaluation standards, indicators, and assessment frameworks for intelligent agents. Different intelligent agent developers or platforms may use their own evaluation methods, resulting in a lack of comparability between evaluation results. This makes it difficult to conduct large-scale and standardized quality control of intelligent agents, affecting the healthy development of the intelligent agent ecosystem. 2. Inefficient and costly evaluation: For subjective and nuanced aspects such as the quality of content generated by intelligent agents, user experience, ethical alignment, or completion of complex tasks, traditional methods often rely heavily on manual evaluation. This is not only time-consuming and labor-intensive, but also costly, and fails to meet the evaluation needs of rapid iteration and large-scale deployment of intelligent agents. 3. The "black box" nature of intelligent agents makes diagnosis difficult: Many current intelligent agents, especially those based on large language models, have a certain "black box" characteristic in their internal decision-making processes and reasoning logic. When an intelligent agent performs poorly or makes errors, it is difficult to perform detailed problem localization, diagnosis, and optimization, which restricts the efficiency of iterative improvement of the intelligent agent; 4. Insufficient Real-World Generalization Validation: Existing benchmarks may not adequately simulate the various situations that agents might encounter in complex real-world application scenarios, especially edge cases, adversarial inputs, or long-duration tasks. There is a gap between the agent's performance on benchmark tests and its generalization ability and robustness after actual deployment, increasing the risk in real-world applications. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for evaluating intelligent agents' network access based on automated tag generation and multi-dimensional assessment, in order to address the aforementioned problems, ensure the quality of the intelligent agent ecosystem, accelerate the commercialization of intelligent agents, and improve the platform's management efficiency for intelligent agents.
[0004] The technical solution of the present invention is as follows: A method for evaluating intelligent agents entering the network based on automated label generation and multi-dimensional assessment includes the following steps: Constructing an intelligent agent profile: Receive the registration information of the intelligent agent to be evaluated, and automatically generate multi-dimensional labels for the intelligent agent based on the registration information using a first large-scale language model guided by structured prompts; collect multi-dimensional evaluation indicators of the intelligent agent during task execution through simulated interaction and testing; Test case generation and evaluation execution: Based on multi-dimensional labels and evaluation metrics, the second large-scale language model is used to automatically generate scenario-based test cases with multi-dimensional labels as key inputs and constraints; the test cases are executed in the test environment, and the behavior trajectory and running data of the agent during the execution of the test cases are recorded; Evaluation Result Analysis and Report Generation: Based on behavioral trajectories and operational data, combined with preset quality standards and scoring algorithms, the intelligent agent is automatically scored and diagnosed for problems, and an evaluation report is generated; Problem diagnosis locates the specific stage of intent understanding, tool selection, data processing, or result generation by analyzing behavioral trajectories and operational data. Network access adaptation and profile update: Based on the scores in the evaluation report, determine whether the agent meets the preset network access specifications; if it does, update the complete agent profile, which includes multi-dimensional tags, evaluation indicators and evaluation report information, to the agent catalog of the multi-agent collaborative platform.
[0005] Furthermore, the first large-scale language model and / or the second large-scale language model are lightweight large-scale language models with a parameter count between 1 billion and 7 billion, and are trained using parameter-efficient fine-tuning techniques and model quantization techniques are employed during deployment.
[0006] Furthermore, the evaluation metrics are collected by deploying a data monitoring agent to capture the agent's operational data in real time, and by using a third large-scale language model as an evaluator to automatically evaluate the agent's output.
[0007] Furthermore, the registration information includes the agent's name, description, skill list, input data type, and output data type.
[0008] Furthermore, the multi-dimensional tags include intelligent agent classification, core functions, application fields, core capabilities, operational entity objects, execution actions, and language emphasis.
[0009] Furthermore, the multi-dimensional evaluation metrics include task success rate, average token consumption, average task time, matching degree between generated content and description, output quality score, stability and error rate, response latency, context understanding ability score, and use case coverage.
[0010] Furthermore, the generation of the scenario-based test cases is based on the functional description of the agent, the application domain, and key actions, and is semantically reasoned and expanded by a second large-scale language model.
[0011] Furthermore, the generation of the multi-dimensional labels includes matching, classifying, and / or expanding the results generated by the large language model with a predefined label vocabulary.
[0012] Furthermore, the evaluation report includes quantitative indicator scores, error details, diagnostic analysis, and improvement suggestions; the agent catalog is a component of the multi-agent collaborative platform, used to support the retrieval, matching, and scheduling of agents.
[0013] This application also includes an agent network entry evaluation system based on automated label generation and multi-dimensional evaluation, using an agent network entry evaluation method based on automated label generation and multi-dimensional evaluation. The system includes: The intelligent agent profile construction module includes: an automated label generation module that uses a lightweight large language model (LLM) as the core analysis engine to receive metadata provided by the intelligent agent to be evaluated during registration. The LLM performs deep semantic analysis and understanding of this text information and guides it to accurately identify and extract information of the corresponding dimensions from the intelligent agent description and skill list through set structured prompts; and an automated evaluation index collection module that performs multiple rounds of simulated calls and tests on the intelligent agent through a simulated interactive environment and an automated test executor, and collects a series of quantitative and qualitative evaluation indexes that reflect the performance of the intelligent agent in real time. The intelligent agent test case generation and evaluation module includes: an automatic test data generation module, which automatically expands or generates targeted, detailed, and specific scenario-based test cases and test data based on the multi-dimensional and structured tags in the constructed intelligent agent profile and the defined evaluation index requirements, utilizing the generation capabilities of a large language model (LLM); an automated evaluation execution module, which executes the automatically generated test cases in an independent simulation environment; and a result analysis and report generation module, which performs in-depth analysis on all runtime data collected in real time during the evaluation execution process, and automatically scores and diagnoses problems of the intelligent agent based on preset quality standards, scoring algorithms, and problem diagnosis rules, and generates a detailed evaluation report. The agent network access adaptation and management module includes: a network access standard judgment module, which determines whether the agent to be evaluated meets the network access requirements based on the final results of the automated evaluation indicators and the platform's preset thresholds and standards for agent "listing" or "network access"; and a profile update and catalog management module, which updates the complete agent profile, including its basic information, automated tags, and evaluation indicators, to the agent catalog database of the multi-agent collaboration platform if the agent passes the evaluation and meets the network access standards. This profile will be used for subsequent user retrieval, intelligent matching, task scheduling, and platform operation analysis, thereby building a high-quality, easily discoverable, and easily invoked agent ecosystem.
[0014] Compared with existing technologies, the advantages of this invention are: 1. Significantly improved standardization and automation of evaluation: By introducing LLM for automatic tag generation and intelligent test case generation, combined with automated execution and evaluation, manual intervention is greatly reduced, making the intelligent agent evaluation process standardized and automated, and improving efficiency several times over; 2. Objective and multi-dimensional evaluation results: Provide a set of evaluation indicators that include performance, efficiency, capability, robustness, security and other dimensions. Combined with automated data collection and LLM evaluation, the evaluation results are more objective and detailed, and comprehensively reflect the overall capabilities of the intelligent agent. 3. Quality and reliability assurance of intelligent agents: Ensure that only intelligent agents that pass rigorous evaluation and meet the specifications can be connected to the network, thus guaranteeing the quality, reliability and security of the entire intelligent agent ecosystem from the source and enhancing user trust; 4. Accelerate the efficiency of agent network access and management: Simplify the agent registration process and accelerate the introduction and iteration of new agents; at the same time, through standardized profiles and evaluation data, greatly improve the platform's ability to accurately manage, efficiently schedule and intelligently recommend agents; 5. Support for efficient evaluation and system operation in resource-constrained environments: By optimizing the LLM in the evaluation system with lightweight fine-tuning, quantitative pruning, and inference, and combined with system-level resource monitoring and management, task priority scheduling, and other strategies, this evaluation system can run efficiently and stably in localized environments with limited resources, such as all-in-one machines. This ensures that the efficiency and resource consumption of the evaluation system itself are within an acceptable range, reducing the cost of evaluation infrastructure. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the method described in this application.
[0016] Figure 2 A flowchart of the method for constructing a profile of the intelligent agent in this application. Detailed Implementation
[0017] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0018] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0019] Please see Figure 1-2 A method for evaluating intelligent agents entering the network based on automated label generation and multi-dimensional assessment, such as... Figure 1 As shown, it includes the following steps: like Figure 2 As shown, the process involves constructing an agent profile: receiving registration information for the agent to be evaluated, including the agent's name, description, skill list, input data type, and output data type. Based on the registration information, a first large-scale language model is used to automatically generate multi-dimensional tags for the agent using structured prompts. These tags include agent classification, core functions, application domains, core capabilities, manipulated entity objects, executed actions, and language emphasis. Generating these multi-dimensional tags involves matching, categorizing, and / or expanding the results generated by the large-scale language model with a predefined tag vocabulary. Through simulated interaction and testing, multi-dimensional evaluation metrics are collected for the agent during task execution. These metrics include task success rate, average token consumption, average task time, matching degree between generated content and description, output quality score, stability and error rate, response latency, context understanding score, and use case coverage. The collection of evaluation metrics involves deploying a data monitoring agent to capture agent runtime data in real time, and using a third large-scale language model as the evaluator to automatically evaluate the agent's output.
[0020] Test Case Generation and Evaluation Execution: Based on multi-dimensional labels and evaluation metrics, the second large-scale language model automatically generates scenario-based test cases using these labels as key inputs and constraints. The generation of scenario-based test cases is based on the agent's functional description, application domain, and key actions, with semantic reasoning and expansion performed by the second large-scale language model. Test cases are executed in the test environment, and the agent's behavioral trajectory and runtime data are recorded during the execution of the test cases. Evaluation Result Analysis and Report Generation: Based on behavioral trajectories and operational data, combined with preset quality standards and scoring algorithms, the intelligent agent is automatically scored and diagnosed for problems, and an evaluation report is generated; Problem diagnosis locates the specific stage of intent understanding, tool selection, data processing, or result generation by analyzing behavioral trajectories and operational data. Network Adaptation and Profile Update: Based on the scores in the evaluation report, it is determined whether the agent meets the preset network access specifications. If it does, the complete agent profile, including multi-dimensional tags, evaluation indicators, and evaluation report information, is updated to the agent catalog of the multi-agent collaborative platform. The evaluation report includes quantitative indicator scores, error details, diagnostic analysis, and improvement suggestions. The agent catalog is a component of the multi-agent collaborative platform, used to support the retrieval, matching, and scheduling of agents.
[0021] The first and / or second large-scale language models are lightweight large-scale language models with between 1 billion and 7 billion parameters, trained using efficient parameter fine-tuning techniques, and deployed using model quantization techniques.
[0022] This application also includes an intelligent agent network access evaluation system based on automated label generation and multi-dimensional assessment, which adopts a modular design and includes: Agent registration interface: Receives basic registration information for the agent to be evaluated, including name, description, URL, provider, version, documentationUrl, authentication, defaultInputModes, defaultOutputModes, and a detailed list of skills (each skill includes name, description, parameters, inputModes, and outputModes). This information is the raw input for building the agent profile.
[0023] Intelligent Agent Profile Construction Module: Automated Tag Generation Module: The core of this module is a lightweight large language model (LLM) with 3-7B parameters, such as models based on the Qwen or Baichuan2 series with fine-tuned instructions. This LLM receives the agent's name, description, and skills text as input, and guides it through semantic understanding and information extraction using carefully designed structured prompts. The prompts not only contain task instructions but also incorporate a few examples to demonstrate how to accurately identify and extract information about the agent on specific evaluation dimensions from complex text, such as: Identify the primary function: Cue words can guide LLM to focus on verb and noun combinations, such as "This agent is mainly used for 'generating reports' and 'analyzing data', please extract its primary function". Identify the Application Domain: Prompt words can guide the LLM to associate industry keywords from the description, such as "The agent description mentions 'financial market analysis' and 'investment decision support', please determine its application domain". Identify key objects and key actions: Cue words can guide the LLM to extract the core objects and behaviors of agent interactions, such as "This agent can process 'Excel spreadsheets' and perform 'pivots', please extract its key objects and actions". LLM maps or generates predefined tags (AgentCategory, PrimaryFunction, ApplicationDomain, CoreCapability, KeyObjects, KeyActions, LanguageFocus) based on the identified functions, domains, capabilities, entities, and actions. InputDataTypes and OutputDataTypes are automatically aggregated by parsing the defaultInputModes and defaultOutputModes fields.
[0024] Automated evaluation metric collection module: This module is deployed in a standalone test environment or via API calls to the agent under test. It includes: Automated test executor: Simulates user behavior or external system calls, sends test requests to the agent under test, and captures the agent's response.
[0025] Data monitoring agent: Monitors various data of the agent in real time when it is executing tasks, such as: start / end timestamps of each task or subtask, used to calculate average task time and response latency; API interaction logs with the underlying LLM (if the agent itself calls the LLM), used to calculate average token consumption; error logs and exception stacks thrown by the agent, used to calculate stability and error rate.
[0026] The LLM-as-a-judge evaluator utilizes another specially tuned lightweight LLM (e.g., a model similar to or smaller than the label generation model) as an automated evaluator. This evaluator receives the agent's output in a simulated task and compares it to a pre-defined gold standard answer, task requirements, or the agent's own description. Guided by a prompt, it automatically scores the output for accuracy, relevance, completeness, logicality, and formatting, thus providing a score for the matching degree between the generated content and the description, and the quality of the output.
[0027] Intelligent agent test case generation and evaluation module: Automatic Test Data Generation Module: This module also relies on an LLM (Reusable Profile Building LLM or Standalone Model), whose input is the agent's multi-dimensional, structured tags and profile information. Based on this information, combined with general knowledge and preset test scenario templates, the LLM automatically expands or generates targeted, detailed, and specific scenario-based test cases and test data. This application emphasizes that the generation of test cases directly and closely depends on the aforementioned automatically generated structured agent profile tags, rather than simply being generated from descriptive text. The LLM uses this structured tag information as the key input and constraint for generation, aiming to simulate real and complex scenarios in which the agent performs key actions (KeyActions) on core entity objects (KeyObjects) within a specific application domain. For example, for an agent with an Application Domain of "Attracting Investment," KeyObjects of "Corporate Financial Reports," and KeyActions of "Analysis," the LLM will generate a test case of "Analyzing the financial reports of a biopharmaceutical company for the past five years and identifying potential risk points," and further generate test data including edge cases such as "missing financial report data" and "non-standard accounting treatment." The goal of LLM is to maximize coverage of the different functional paths and edge cases represented by these tag combinations in order to improve use case coverage.
[0028] Automated evaluation execution module: In the test environment, this module calls the agent under test sequentially or in parallel according to automatically generated test cases to simulate the complete task flow. During this process, it records in detail the input, output, internal state changes (if observable), tool call sequence, and user intervention points (if the evaluation includes human-machine collaborative simulation) for each call of the agent.
[0029] The results analysis and report generation module performs in-depth analysis of all operational data collected in real time during the evaluation process. Combining preset quality standards, scoring algorithms, and problem diagnosis rules, it automatically scores and diagnoses problems in the agent, generating a detailed evaluation report. This application emphasizes that the specific technical means of problem diagnosis lies in utilizing recorded behavioral trajectories (including tool call sequences, intermediate steps, and decision paths) and detailed operational data (including error stack information, API call failure logs, resource peaks, response timeouts, etc.), combined with preset diagnostic rules or lightweight analysis models (such as rule-based expert systems or small classification models), to pinpoint the specific stages where the fault occurs, such as intent understanding, tool selection, data processing, and result generation. For example, if the trajectory shows that the agent selected the wrong tool in the "intent understanding" stage, or experienced resource exhaustion in the "data processing" stage, the report will clearly indicate the problematic stage and relevant evidence. The report content not only includes quantitative scores and error rates for various indicators, but also provides accurate problem diagnosis and targeted improvement suggestions.
[0030] Intelligent Agent Network Access Adaptation and Management Module: Network Access Standard Judgment Module: This module defines a configurable network access rule engine. Based on the final scores of various evaluation indicators in the evaluation report, it compares them with the platform's preset minimum network access thresholds (e.g., task success rate must be greater than 90%, output quality score higher than 4 points, no security vulnerabilities, etc.) to automatically determine whether the intelligent agent meets the "listing" or "network access" standards.
[0031] Profile Update and Directory Management Module: Once an agent passes the network access evaluation, this module will trigger a data update operation. It will update and synchronize a complete agent profile, including basic agent information, automated tags generated by LLM (such as AgentCategory, PrimaryFunction, etc.), and automated evaluation metrics (such as task success rate, average token consumption, etc.), to the agent directory database of the multi-agent collaborative platform. This complete profile will serve as an important basis for platform user retrieval, intelligent recommendation, task scheduling, and platform operation analysis.
[0032] System model training process implementation details: LLM Fine-Tuning: All LLMs used in this application (for label generation, LLM-as-a-judge evaluation, test case generation, etc.) will employ Parameter-Efficient Fine-Tuning (PEFT) techniques, such as LoRA (Low-Rank Adaptation) or Q-LoRA. This method can freeze the backbone parameters of a pre-trained large model and train only a small number of newly added adapter parameters, significantly reducing the computational resources and memory usage required for training, thus adapting to the limited hardware capabilities of integrated machines.
[0033] Label generation training set: Construct a high-quality dataset of <text, label JSON> pairs of agent description text and corresponding multi-dimensional labels. Data sources include publicly available agent registration information, manually annotated agent descriptions, and labels generated through expert alignment.
[0034] Evaluation training set: contains agent test case inputs, agent outputs, and corresponding quality evaluation results (input, output, evaluation score) pairs, used to train the LLM-as-a-judge model to learn human evaluation standards.
[0035] Test case generation training set: A dataset containing agent profile labels and corresponding high-quality test cases in <profile JSON, test case text> pairs, used to train the LLM to learn how to generate meaningful test cases based on agent characteristics.
[0036] Data preprocessing: Perform necessary cleaning, word segmentation, stop word removal, normalization, and other operations on the training data, and convert it into an input format acceptable to the model (such as a token ID sequence).
[0037] Loss function and optimizer: The cross-entropy loss function, suitable for text generation and classification tasks, is adopted. The optimizer selected is AdamW, etc., combined with adaptive learning rate adjustment algorithms (such as Adafactor, Adagrad-Delta) to accelerate model convergence while avoiding gradient explosion or vanishing.
[0038] Distributed training (optional): If the all-in-one machine supports multiple GPUs, a distributed training strategy using data parallelism (such as PyTorch's DistributedDataParallel, DDP) can be adopted to further accelerate the training process.
[0039] Mixed precision training: During training, mixed precision training of FP16 or FP8 (if supported by hardware and software libraries) is used to reduce memory usage and accelerate computation.
[0040] Deployment and optimization: Model compression: Quantization is performed on the trained LLM model, for example, converting parameters from FP16 / FP32 to INT8, which significantly reduces the model size and computational cost. Simultaneously, model pruning can be considered to remove redundant connections and neurons, further compressing the model and improving inference efficiency.
[0041] High-performance inference engine: When deploying models on an all-in-one machine, high-performance inference frameworks such as TensorRT, ONNX Runtime, or OpenVINO are utilized. These frameworks can perform model graph optimization, kernel fusion, and efficient memory management for specific hardware (such as NVIDIA GPUs), maximizing inference throughput and reducing latency.
[0042] Dynamic Batching: During inference, the input batch size is dynamically adjusted according to the real-time request load, effectively utilizing the parallel computing capabilities of the all-in-one GPU to improve inference throughput.
[0043] Heterogeneous hardware adaptation: When designing the evaluation system, the collaborative work of different computing units (CPU, GPU, NPU) is considered, and different models or computing tasks are assigned to the most suitable hardware for execution.
[0044] Containerized deployment: The entire evaluation system (including LLM service, test executor, and data analysis module) is packaged into lightweight Docker containers. This facilitates rapid deployment on appliance platforms, environment isolation, version management, and horizontal scaling, improving system robustness and maintainability.
[0045] System-level resource monitoring and management: The evaluation system itself integrates an adaptive resource monitoring and management mechanism to track the utilization of system resources such as CPU, memory, and GPU in real time. Through task priority scheduling and dynamic resource allocation strategies, it ensures that evaluation tasks run efficiently with limited resources, avoiding resource exhaustion or mutual interference, thereby ensuring that the efficiency and resource consumption of the evaluation system itself are always within acceptable limits. For example, when GPU resources are scarce, the system can automatically reduce the concurrency of non-critical LLM evaluation tasks or prioritize high-priority test case generation tasks.
[0046] The embodiments described above merely illustrate specific implementation methods of this application, and while the descriptions are detailed and specific, they should not be construed as limiting the scope of protection of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the technical solution of this application, and these modifications and improvements all fall within the scope of protection of this application.
Claims
1. A method for evaluating the network entry of intelligent agents based on automated label generation and multi-dimensional assessment, characterized in that, Includes the following steps: Constructing an intelligent agent profile: Receive the registration information of the intelligent agent to be evaluated, and automatically generate multi-dimensional labels for the intelligent agent based on the registration information using a first large-scale language model guided by structured prompts; collect multi-dimensional evaluation indicators of the intelligent agent during task execution through simulated interaction and testing; Test case generation and evaluation execution: Based on multi-dimensional labels and evaluation metrics, the second large-scale language model is used to automatically generate scenario-based test cases with multi-dimensional labels as key inputs and constraints. Execute test cases in the test environment and record the behavior trajectory and running data of the agent during the execution of the test cases; Evaluation Result Analysis and Report Generation: Based on behavioral trajectories and operational data, combined with preset quality standards and scoring algorithms, the intelligent agent is automatically scored and diagnosed for problems, and an evaluation report is generated; Problem diagnosis locates the specific stage of intent understanding, tool selection, data processing, or result generation by analyzing behavioral trajectories and operational data. Network access adaptation and profile update: Based on the scores in the evaluation report, determine whether the agent meets the preset network access specifications; if it does, update the complete agent profile, which includes multi-dimensional tags, evaluation indicators and evaluation report information, to the agent catalog of the multi-agent collaborative platform.
2. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The first large-scale language model and / or the second large-scale language model are lightweight large-scale language models with between 1 billion and 7 billion parameters, trained using efficient parameter fine-tuning techniques, and deployed using model quantization techniques.
3. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The evaluation metrics are collected by deploying a data monitoring agent to capture the agent's operational data in real time, and by using a third large-scale language model as an evaluator to automatically evaluate the agent's output.
4. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The registration information includes the agent's name, description, skill list, input data type, and output data type.
5. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The multi-dimensional tags include agent classification, core functions, application areas, core capabilities, manipulated entity objects, executed actions, and language emphasis.
6. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The multi-dimensional evaluation metrics include task success rate, average token consumption, average task time, matching degree between generated content and description, output quality score, stability and error rate, response latency, context understanding ability score, and use case coverage.
7. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The generation of the scenario-based test cases is based on the functional description of the agent, the application domain, and key actions, and is performed by semantic reasoning and expansion by a second large-scale language model.
8. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The generation of the multi-dimensional labels includes matching, classifying, and / or expanding the results generated by the large language model with a predefined label vocabulary.
9. The method for evaluating intelligent agents entering the network based on automated tag generation and multi-dimensional assessment according to claim 1, characterized in that, The evaluation report includes quantitative indicator scores, error details, diagnostic analysis, and improvement suggestions; The agent catalog is a component of a multi-agent collaborative platform, used to support the retrieval, matching, and scheduling of agents.
10. A smart agent network access evaluation system based on automated label generation and multi-dimensional assessment, characterized in that, Using the intelligent agent network access evaluation method based on automated label generation and multi-dimensional evaluation as described in any one of claims 1-9, the system includes: The intelligent agent profile construction module includes: an automated label generation module that uses a lightweight large language model (LLM) as the core analysis engine to receive metadata provided by the intelligent agent to be evaluated during registration. The LLM performs deep semantic analysis and understanding of this text information and guides it to accurately identify and extract information of the corresponding dimensions from the intelligent agent description and skill list through set structured prompts; and an automated evaluation index collection module that performs multiple rounds of simulated calls and tests on the intelligent agent through a simulated interactive environment and an automated test executor, and collects a series of quantitative and qualitative evaluation indexes that reflect the performance of the intelligent agent in real time. The intelligent agent test case generation and evaluation module includes: an automatic test data generation module, which automatically expands or generates targeted, detailed, and specific scenario-based test cases and test data based on the multi-dimensional and structured tags in the constructed intelligent agent profile and the defined evaluation index requirements, utilizing the generation capabilities of a large language model (LLM); an automated evaluation execution module, which executes the automatically generated test cases in an independent simulation environment; and a result analysis and report generation module, which performs in-depth analysis on all runtime data collected in real time during the evaluation execution process, and automatically scores and diagnoses problems of the intelligent agent based on preset quality standards, scoring algorithms, and problem diagnosis rules, and generates a detailed evaluation report. The agent network access adaptation and management module includes: a network access standard judgment module, which determines whether the agent to be evaluated meets the network access requirements based on the final results of the automated evaluation indicators and the platform's preset thresholds and standards for agent "listing" or "network access"; and a profile update and directory management module, which updates the complete agent profile, including its basic information, automated tags, and evaluation indicators, to the agent directory database of the multi-agent collaboration platform if the agent passes the evaluation and meets the network access standards. This profile will be used for subsequent user retrieval, intelligent matching, task scheduling, and platform operation analysis, thereby building a high-quality, easily discoverable, and easily invoked agent ecosystem.
Citation Information
Cited By
Inductive process mining optimization method and system based on multi-agent cooperation
CN121436623A
Inductive process mining optimization method and system based on multi-agent cooperation
CN121436623B
End-to-end AI agent evaluation method, system, equipment and medium
CN121900800A
Intelligent agent self-error correction method and system based on track comparison and positioning
CN122173330A
An agent self-correction method and system based on trajectory comparison and positioning
CN122173330B