Multi-agent system evaluation method, device and system, electronic equipment and medium
By receiving and analyzing data from multi-agent systems under different communication states, and calculating multi-dimensional quantitative indicators, the limitations and lack of transparency of existing evaluation technologies are resolved, enabling reliable evaluation and certification of multi-agent systems and improving the accuracy and reliability of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUA CHUAN INTERNATIONAL HOLDINGS GROUP CO LTD
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-24
AI Technical Summary
Existing multi-agent system evaluation technologies suffer from problems such as one-sided evaluation dimensions, non-standard test scenarios, and opaque evaluation results. These issues prevent the effective measurement of collective intelligence and hinder the healthy development and safe application of multi-agent systems.
By receiving operational data from multi-agent systems under normal and shielded communication conditions, the system calculates collective output, individual output, and collaborative efficiency indicators. Combined with auxiliary indicators such as fairness, compliance with norms, and adaptive recovery rate, it generates multi-dimensional quantitative indicators, providing a standardized, quantitative, and interpretable evaluation method.
It enables standardized evaluation of the collective, social, and systemic aspects of multi-agent systems, improves the reliability and accuracy of evaluation, provides visualized certification reports and improvement suggestions, and enhances the trustworthiness of the system.
Smart Images

Figure CN121920654A_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of artificial intelligence technology, specifically relating to technologies such as intelligent agent performance evaluation and large-scale language models, and in particular to a method and apparatus, system, electronic device and computer-readable storage medium for evaluating multi-agent systems. Background Technology
[0002] With the exponential growth of computing power and continuous breakthroughs in algorithms, multi-agent systems have rapidly evolved from theoretical models in academia into core technological engines driving change in numerous key fields. In the transportation sector, fleets of dozens or even hundreds of autonomous vehicles are expected to solve urban congestion and improve logistics efficiency; in industrial automation, hundreds or thousands of warehouse robots work collaboratively in massive warehouses, forming the logistics backbone of modern e-commerce; in defense and security, the collaborative combat capabilities of drone swarms are seen as key to disrupting the rules of future battlefields; and in digital entertainment, groups capable of exhibiting complex social behavior and collective intelligence are becoming the cornerstone of building immersive virtual worlds.
[0003] However, despite the promising prospects of MAS (Multi-Agent Systems), the technical means for testing, evaluation, and certification lag far behind the pace of its application development, creating a widening evaluation deficit. This deficit not only hinders the healthy development of the technology but also sows huge, unquantified risks in key application areas. Summary of the Invention
[0004] This disclosure provides a method and apparatus for evaluating multi-agent systems, a system, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect, a method for evaluating a multi-agent system is provided. The method includes: receiving first operational data and second operational data; the first operational data is information interaction or environmental perception sharing data between agents in the multi-agent system under test acquired after the system is connected to a pre-configured test environment; the second operational data is the state of each agent or local environmental decision data obtained after adjusting or disabling the interaction capabilities between agents in the multi-agent system under test in the test environment; calculating the collective output value of the multi-agent system under test based on the first operational data; calculating the individual output value of each agent based on the second operational data, and determining the sum of individual output values; calculating a collaboration index value of a collaboration efficiency index based on the collective output value and the sum of individual output values; calculating an auxiliary index value based on the first operational data; the auxiliary index value includes at least one of a fairness index, a norm compliance index, and an adaptive recovery rate index; and generating multi-dimensional quantitative index values of the multi-agent system under test based on the collaboration index value and the auxiliary index value.
[0006] According to a second aspect, a multi-agent system evaluation apparatus is provided, comprising: a receiving unit configured to receive first operational data and second operational data, wherein the first operational data is information interaction or environmental perception sharing data between agents in the multi-agent system under test acquired after the multi-agent system under test is connected to a pre-configured test environment; and the second operational data is the state of each agent or local environmental decision data of each agent obtained after adjusting or disabling the interaction capabilities between agents in the multi-agent system under test in the test environment; and a collective value calculation unit configured to calculate the multi-agent system under test based on the first operational data. The system comprises: a collective output value; an individual value calculation unit configured to calculate the individual output value of each agent based on the second operating data, and determine the sum of the individual output values; a collaborative calculation unit configured to calculate the collaborative efficiency index value based on the collective output value and the sum of the individual output values; an auxiliary calculation unit configured to calculate auxiliary index values based on the first operating data; the auxiliary index values include at least one of the following: fairness index value, norm compliance index value, and adaptive recovery rate index value; and a generation unit configured to generate multi-dimensional quantitative index values of the multi-agent system under test based on the collaborative index values and the auxiliary index values.
[0007] According to the third aspect, a multi-agent system evaluation system is provided, comprising: an access module configured to receive a multi-agent system under test (MAS) packaged into standardized deployment units; an orchestration module configured to dynamically create isolated test environments and deploy the MAS and a standardized virtual simulation platform within these test environments; a test control module configured to control the virtual simulation platform to execute a test process, the test process including: enabling inter-agent interaction in the MAS to execute collaborative operation tests and collecting first operation data in a first time period; and disabling inter-agent interaction in the MAS to execute independent operation tests and collecting second operation data in a second time period; a core computing engine configured to receive the first and second operation data, and based on the first and second operation data, execute the method of the first aspect to calculate multi-dimensional quantitative indicators, including collaborative efficiency; and a report generation module configured to generate a visualized certification report based on the multi-dimensional quantitative indicators.
[0008] According to a fourth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.
[0009] According to a fifth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any implementation of the first aspect.
[0010] The multi-agent system evaluation method and apparatus provided in the embodiments of this disclosure firstly receive first operating data and second operating data. The first operating data is information interaction or environmental perception sharing data between agents in the multi-agent system under test obtained after the multi-agent system under test is connected to a pre-configured test environment. The second operating data is the state of each agent or local environmental decision data of each agent obtained after adjusting or disabling the interaction capabilities between agents in the multi-agent system under test in the test environment. Secondly, based on the first operating data, the collective output value of the multi-agent system under test is calculated. Thirdly, based on the second operating data, the individual output value of each agent is calculated, and the sum of individual output values is determined. Next, based on the collective output value and the sum of individual output values, a collaboration index value of the collaboration efficiency index is calculated. Then, based on the first operating data, an auxiliary index value is calculated. The auxiliary index value includes at least one of a fairness index value, a standard compliance index value, and an adaptive recovery rate index value. Finally, based on the collaboration index value and the auxiliary index value, a multi-dimensional quantitative index value of the multi-agent system under test is generated. Therefore, it is possible to conduct standardized, quantitative, and interpretable evaluation and certification of multi-agent systems from the dimensions of collectivity, sociality, and systemicity, thereby improving the reliability and accuracy of multi-agent evaluation.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0013] Figure 1 This is a flowchart of an embodiment of the multi-agent system evaluation method according to the present disclosure;
[0014] Figure 2 This is a schematic diagram illustrating the adaptive recovery rate calculation process in this disclosure;
[0015] Figure 3 This is a schematic diagram of a structure of an embodiment of the multi-agent system evaluation device according to the present disclosure;
[0016] Figure 4 This is a structural diagram of a multi-agent system evaluation system according to the present disclosure;
[0017] Figure 5This is a block diagram of an electronic device used to implement the multi-agent system evaluation method of the embodiments of this disclosure. Detailed Implementation
[0018] Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprises" shall be understood to include the stated elements or components without excluding other elements or other components.
[0019] The technical solutions of this disclosure are illustrated below through specific embodiments. It should be understood that one or more steps mentioned in this disclosure do not preclude the existence of other methods and steps before or after the combined steps, or that other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Unless otherwise stated, the numbering of each method step is only for the purpose of identifying each method step, and not to limit the order of each method or to limit the scope of implementation of this disclosure. Changes or adjustments to their relative relationships, without substantial changes to the technical content, can also be considered as within the scope of implementation of this disclosure.
[0020] The raw materials and instruments used in the examples are not subject to any specific restrictions on their source; they can be purchased from the market or prepared according to conventional methods known to those skilled in the art.
[0021] Current methods for evaluating intelligent agent systems in the background technology mainly expose the following three fundamental defects:
[0022] First, there is an individualism fallacy in evaluation metrics. Most existing evaluation frameworks are based on the core idea of evaluating individual intelligence, simply viewing a multi-agent system as a collection of independent agents. Therefore, evaluation metrics are often linear sums of individual capabilities. For example, in the evaluation of autonomous driving fleets, the industry generally focuses on the perception accuracy of individual vehicles, the success rate of planning algorithms, and end-to-end mileage. While these metrics are important, they completely fail to answer a deeper and more crucial question: when these "excellent" individual vehicles merge into real, complex mixed traffic flows, can they, as a whole, exhibit efficient, safe, and even courteous collective behavior? A fleet where all vehicles have an individual safety score of 99% does not mean that the fleet as a whole has an accident probability of 1% to the power of N. On the contrary, negative synergistic effects between individuals, such as phantom traffic jams, are caused by multiple excellent individuals pursuing local optima. Current technology lacks a set of metrics that can directly measure collective and emergent characteristics such as collaborative gain, system resilience, and fairness in resource allocation.
[0023] Secondly, there is the challenge of sterile laboratory conditions in the evaluation environment. Current MAS testing is mostly conducted in highly simplified and idealized simulation environments, or in closed, controlled physical test sites. These "sterile laboratory" environments, for the sake of testing convenience, often filter out the most fatal source of complexity in the real world: social dynamics. For example, in testing warehouse robots, the environment might only contain static shelves and predictable robot companions. However, in a real warehouse, there are workers with unpredictable walking routes, suddenly falling packages, and robots developed by other companies that follow different behavioral logics. Existing testing environments cannot effectively simulate this heterogeneous, uncertain scenario that requires social negotiation. Therefore, a robot swarm that performs perfectly in a "sterile laboratory" may experience a catastrophic performance degradation when deployed in a real environment.
[0024] Third, there is a black box problem in the evaluation results. Even if some tests can provide a comprehensive final score, such as the total mission completion time, this single number masks the complex processes behind it. We cannot know from this result whether the system's efficiency stems from a truly effective collaborative understanding, or simply from good luck, or from a single super-individual compensating for the inefficiencies of all other members. This lack of interpretability is unacceptable in critical areas requiring high reliability and trustworthiness. For example, a commander of a military drone swarm needs to know whether his swarm is a disciplined, well-formed Roman legion or a band of rogue pilots relying on a few elite pilots. These two internal collaborative models are drastically different in their robustness to various battlefield disturbances. Existing technology cannot provide a diagnostic tool to dissect the black box of the MAS and gain insight into its internal collaborative patterns, normative structures, and potential vulnerabilities.
[0025] In summary, a significant technological gap exists in the field of MAS (Multi-Agent Systems) evaluation: while AI groups are capable of complex interactions, there is no reliable standard to measure their collective intelligence. This situation not only makes it difficult for users to effectively select from different MAS products on the market, but also prevents regulatory agencies from developing scientific and reasonable access standards for this powerful technology. This lack of evaluation capability has become a core bottleneck restricting the entire multi-agent systems industry from technology demonstrations to large-scale reliable applications.
[0026] To address the fundamental technical problems in existing technologies, such as one-sided evaluation dimensions for multi-agent systems, non-standard testing scenarios, and opaque evaluation results, this disclosure provides a method for evaluating multi-agent systems. Figure 1 A flowchart 100 is shown as an embodiment of a multi-agent system evaluation method according to the present disclosure, the multi-agent system evaluation method comprising the following steps:
[0027] Step 101: Receive the first running data and the second running data.
[0028] In this embodiment, the first operational data is the information interaction or environmental perception sharing data between the agents in the multi-agent system under test (MAS) obtained after the MAS connects to a pre-configured test environment. The second operational data is the individual state or local environmental decision data of each agent obtained after adjusting or disabling the interaction capabilities between the agents in the MAS in the test environment. It should be noted that the first and second operational data can be real-time data or historically stored data. The test environment is the environment in which the MAS is located; the test environment can be a virtual simulation platform, a physical test field, or a hybrid virtual-physical environment. The second operational data includes the individual state or local environmental decision data of each agent obtained after adjusting (e.g., reducing communication bandwidth or increasing communication latency) the interaction capabilities between the agents in the MAS in the test environment, or the individual state or local environmental decision data of each agent obtained after disabling the interaction capabilities between the agents in the MAS in the test environment.
[0029] In this embodiment, when the multi-agent system under test (MAS) performs a certain task, the system collects first operational data when maintaining normal communication and collaboration. This data consists of information exchanged between the agents and shared environmental perception data. At this time, each agent in the MAS is responsible for its corresponding part of the task. Subsequently, in the same simulation scenario, the interaction channels between all agents are forcibly cut off, leaving only their independent perception-decision loops. The system then collects second operational data, which consists of independent decision data made by each agent based solely on its own state and local environment for its respective sub-task. By comparing the first and second operational data, the impact of information interaction on the overall system behavior and individual decisions can be quantitatively evaluated.
[0030] Step 102: Calculate the collective output value of the multi-agent system under test based on the first running data.
[0031] In this embodiment, the collective output value is the data produced by all agents in the multi-agent system under test in the test environment (such as the number of tasks completed, the coverage area, etc.).
[0032] In this embodiment, based on the first operating data, the total number of effective tasks, total resource utilization, or cumulative value of the objective function completed by all agents in the multi-agent system under test within the test environment and observation period are statistically analyzed and used as the collective output value of the multi-agent system under test.
[0033] Step 103: Based on the second running data, calculate the individual output value of each agent and determine the sum of the individual output values.
[0034] In this embodiment, the individual output value is the data (such as the number of tasks completed, the coverage area, etc.) produced by each agent in the test environment after the interaction ability between each agent in the multi-agent system under test is cut off. The sum of the individual output values is the sum of the individual output values of the multi-agent system under test.
[0035] In this embodiment, based on the second running data, the effective task volume, resource utilization volume, or individual objective function value independently completed by each agent in the same cycle are calculated. The calculated effective task volume, resource utilization volume, or individual objective function value is used as the individual output value. Then, the individual output values of all agents in the multi-agent system under test are added together to obtain the sum of the individual output values of each agent.
[0036] Step 104: Calculate the collaboration index value of the collaboration efficiency index based on the sum of the collective output value and the individual output value.
[0037] In this embodiment, the collaboration efficiency index is an indicator for evaluating the collaboration efficiency between agents in the multi-agent system under test, and the collaboration index value is the value of the collaboration efficiency index.
[0038] In this embodiment, the collaboration index value is calculated by comparing the collective output value with the sum of all individual output values. Specifically, the system counts the total collective output completed by all members during the collaboration period, and simultaneously adds up the individual output value completed independently by each member to obtain the sum of individual output values. The collective output value is then divided by the sum of individual output values, and the resulting ratio is the collaboration index value. The collaboration efficiency index is used to quantify the degree of efficiency improvement of team collaboration relative to individual independent work.
[0039] Step 105: Calculate auxiliary indicator values based on the first running data.
[0040] In this embodiment, the auxiliary index values include at least one of the following: fairness index value, norm compliance index value, and adaptive recovery rate index value. The fairness index value is a fairness indicator (such as balance, equilibrium, coordination, symmetry, uniformity, etc.), used to characterize the degree of balance in resource allocation within the multi-agent system under test. For example, the fairness index is the Gini coefficient based on waiting time or the Jain fairness index based on cumulative returns. The norm compliance index value is a norm compliance indicator (such as compliance, conformity, adherence, achievement rate, consistency), used to characterize the degree of norm compliance by agents in the multi-agent system under test. The adaptive recovery rate index value is an adaptive recovery rate indicator (such as robustness, stability, reliability, fault tolerance, disturbance resistance), used to characterize the degree of disturbance susceptibility of the multi-agent system under test. For example, the adaptive recovery rate index is calculated based on the performance recovery magnitude or the integral area of performance loss.
[0041] In this embodiment, the auxiliary indicator value can be any one of the fairness indicator value, the standard compliance indicator value, and the adaptive recovery rate indicator value. The auxiliary indicator value can also be a value obtained by calculating multiple of the fairness indicator value, the standard compliance indicator value, and the adaptive recovery rate indicator value, such as by weighted summation of the fairness indicator value, the standard compliance indicator value, and the adaptive recovery rate indicator value to obtain the auxiliary indicator value.
[0042] In this embodiment, the fairness index can be obtained through the calculation of the Jain fairness index based on resource utility or reward. This Jain fairness index measures the amount of resources or task rewards obtained by the agent. Furthermore, the Jain index, commonly used in computer networks and communications, is adopted. This Jain index is more sensitive to multi-agent systems with a small sample size (e.g., 4-10 robots) and has better normalization effects. Optionally, the fairness index can also employ Shannon entropy. The total movement distance or total energy consumption of each agent in the system under test is input into the Shannon entropy formula to calculate the proportion of each agent's workload to the total workload. This Shannon entropy formula introduces an information theory perspective and is suitable for scenarios that are extremely sensitive to energy consumption balance.
[0043] In this embodiment, the compliance index can assess whether an agent is on the verge of violating regulations for an extended period. The steps for calculating the compliance index are as follows: A virtual repulsive potential field is constructed around the restricted area or the agent's peers; the potential field strength increases exponentially with decreasing distance. At each point on the trajectory, the potential field strength of the agent in the system under test is calculated. The maximum permissible potential field threshold for the average potential field strength of all agents in the system under test is calculated. This maximum permissible potential field threshold is then subtracted from 1 to obtain the compliance index. Thus, the evaluation of the agent system can be shifted from post-hoc statistics to process risk measurement.
[0044] In this embodiment, in certain battlefield or high-frequency trading scenarios, the speed of recovery is more important than the degree of recovery. The adaptive recovery rate metric is measured as the time required for performance to return to a specific percentage (e.g., 90%) of the baseline after a disturbance occurs. The specific implementation steps are: marking the disturbance injection time t. event Set the recovery threshold P. thresh = P baseline \times 90%. In t>t event Find the first one that satisfies the data. And the time t remains stable thereafter recovered The calculation of the adaptive recovery rate index is shown in Equation (1).
[0045] ARRtime=1+log(t recovered -t event(1) In equation (1), ARRtime represents the adaptive recovery rate index value. The adaptive recovery rate index value uses a logarithmic or reciprocal function so that the shorter the time, the higher the score. The adaptive recovery rate index value covers the evaluation of the agility of the agent in the tested agent system.
[0046] Optionally, after acquiring the first operational data of the multi-agent system under test, the resource consumption, rule violation count, and timestamps of the occurrence and recovery of interference events for each agent at different times are extracted. Based on this, the following indicators are calculated: Fairness index value—using the Gini coefficient to measure the degree of balance in resource allocation, with a value closer to 0 indicating greater fairness; Rule compliance index value—reflecting the system's compliance level by the proportion of runtime without violations to the total runtime; Adaptive recovery rate index value—statistically calculating the success rate of the system recovering to a steady state within a specified time after external interference occurs, used to quantify its anti-interference and self-recovery capabilities; Finally, auxiliary index values are output to provide a quantitative basis for system performance evaluation.
[0047] Step 106: Generate multi-dimensional quantitative index values for the multi-agent system under test based on the collaborative index value and the auxiliary index value.
[0048] In this embodiment, the multidimensional quantitative index value is the value of the multidimensional quantitative index, which is an index that measures the multi-agent system under test in multiple dimensions. The multidimensional quantitative index is also known as C-IQ (Collective Intelligence Quotient), and therefore the multidimensional quantitative index value is also referred to as the C-IQ total score in this disclosure.
[0049] In this embodiment, the collaboration index value and the auxiliary index value are first normalized, and then weighted and fused according to a preset weight vector to obtain the "collaboration-auxiliary comprehensive score". Then, this score is concatenated with system layer performance, robustness, communication efficiency and other dimensional indicators to form a feature vector. After dimensionality reduction by principal component analysis, it is mapped to the [0, 1] interval. Finally, a multi-dimensional quantitative index value that reflects the level of collaboration, auxiliary contribution and overall effectiveness is output, so as to realize a unified quantitative evaluation of the collaboration and auxiliary capabilities of multi-agent system.
[0050] The multi-agent system evaluation method provided in this disclosure first receives first operational data and second operational data. The first operational data is information interaction or environmental perception sharing data between agents in the multi-agent system under test, obtained after the system is connected to a pre-configured test environment. The second operational data is the state of each agent or local environmental decision data obtained after adjusting or disabling the interaction capabilities between agents in the multi-agent system under test in the test environment. Second, based on the first operational data, the collective output value of the multi-agent system under test is calculated. Third, based on the second operational data, the individual output value of each agent is calculated, and the sum of individual output values is determined. Fourth, based on the collective output value and the sum of individual output values, a collaboration index value of the collaboration efficiency index is calculated. Then, based on the first operational data, an auxiliary index value is calculated. Finally, based on the collaboration index value and the auxiliary index value, a multi-dimensional quantitative index value of the multi-agent system under test is generated. Therefore, a standardized, quantifiable, and interpretable evaluation and certification of multi-agent systems can be performed from the dimensions of collectivity, sociality, and systemicity, improving the reliability and accuracy of multi-agent evaluation.
[0051] In one embodiment of this disclosure, the multi-agent evaluation method further includes: generating a visualized authentication report based on multi-dimensional quantitative index values.
[0052] In this embodiment, the certification report is not only a score, but also a diagnostic report for the multi-agent system under test. The certification report includes: multi-dimensional quantitative index values and multi-dimensional quantitative index value derivative graphs (such as radar charts of various indicators in the multi-dimensional quantitative index corresponding to the multi-dimensional quantitative index values); detailed curves of the changes of various indicators of the multi-agent system under test (such as throughput and conflict rate) over time; automatically captured and visualized key event replay information of the "highlight moments" (such as a successful complex collaboration) and the darkest moments (such as a systemic collapse) in the testing process; based on the score shortcomings, an expert system or large language model automatically generates targeted improvement suggestions, such as the improvement suggestion: Your system scores low on GWT (Global Workspace Theory), and it is recommended to optimize the resource queuing algorithm.
[0053] In some embodiments of this disclosure, the above-mentioned calculation of the collaboration index value based on the sum of collective output value and individual output value includes: subtracting the sum of individual output value from the collective output value to obtain the net gain value; determining the collaboration index value based on the ratio of the net gain value to the sum of individual output value; wherein, if the sum of individual output value is zero and the collective output value is greater than zero, the collaboration index value is set to a preset saturation value.
[0054] In this optional implementation, when the sum of individual output values is zero and the collective output value is greater than zero, the collaboration index value is set to a preset saturation value, which can effectively characterize the system's strong dependence on collaboration.
[0055] Alternatively, the net gain can be obtained by subtracting the sum of individual outputs from the collective output value, and the cooperation index value SE can be obtained by dividing the net gain by the sum of individual outputs. Specifically, the collective output value is... The individual output value of N (N>1) agents operating independently in the system under test. The total output of each individual is calculated as shown in equation (2).
[0056] (2)
[0057] The net gain from computational collaboration is shown in equation (3).
[0058] (3)
[0059] The calculated collaboration index value SE is shown in equation (4).
[0060] (4)
[0061] In equation (4), in order to prevent the denominator from being zero, when the sum of individual outputs is zero, if the collective output value is greater than zero, the cooperation index value is directly set to the preset saturation value (such as 1), otherwise the cooperation index value is set to 0.
[0062] In this alternative implementation, the difficulty in calculating the collaboration metric (SE) lies in defining a fair sum of collective and individual output values for complex tasks. This disclosure proposes a task graph decomposition method: decomposing complex tasks into a series of dependent sub-task graphs. The sum of collective and individual output values is defined as the weighted sum of completed sub-task nodes, with the weight determined by the criticality of the sub-task in the graph.
[0063] The method for calculating the collaboration index value provided in the embodiments of this disclosure calculates the net gain value by summing the collective output value and the individual output value, and obtains the collaboration index value based on the ratio of the net gain value to the sum of the individual output values, thus providing a reliable means of obtaining the collaboration index value.
[0064] In some embodiments of this disclosure, the above-mentioned auxiliary indicator value is a fairness indicator value. The above-mentioned calculation of the auxiliary indicator value based on the first running data includes: extracting the resource request events of each agent for the shared resources and the corresponding waiting time from the first running data; calculating a statistical coefficient reflecting the dispersion of resource allocation based on the waiting time distribution of all agents, and using the statistical coefficient as a fairness indicator value.
[0065] In this optional implementation, each intelligent agent Waiting duration sequence within the task cycle Calculate the total waiting time for each agent. For the total waiting time sequence Sort the data in ascending order and calculate the total waiting time. Based on the total waiting time and the cumulative waiting time ratio for each sorting position, the Gini coefficient GET is calculated using the standard Gini coefficient calculation formula as shown in equation (5), and the Gini coefficient GET is used as a statistical coefficient:
[0066] (5)
[0067] In this optional implementation, the execution entity includes an event detector, which can automatically identify key events such as requesting resources, starting to wait, and acquiring resources from the trajectory data, so as to accurately calculate the total waiting time. The executing entity identifies waiting events for mutually exclusive shared resources for each agent from the initial operational data using an event detector and extracts the waiting duration sequence. Based on the waiting duration sequence, a Lorenz curve representing the resource allocation distribution is constructed. Finally, based on the Lorenz curve and the absolute fairness line, a fairness index value is calculated. This provides a reliable method for obtaining the fairness index value.
[0068] Optionally, the auxiliary indicator value is a fairness indicator value. Based on the first operating data, the calculation of the auxiliary indicator value includes: extracting the cumulative revenue value or resource possession obtained by each agent within a preset test period from the first operating data; calculating the ratio between the square of the sum of the cumulative revenue values of all agents and the total number of agents multiplied by the sum of the squares of the cumulative revenue values of each agent; and using the ratio as a fairness indicator value to characterize the uniformity of positive incentive allocation within the system.
[0069] Optionally, the auxiliary indicator value is a fairness indicator value. Based on the first operating data, the calculation of the auxiliary indicator value includes: extracting the energy consumption data or movement distance data of each agent from the first operating data as load data; calculating the proportion of the load data of each agent to the total load data of all agents to obtain the load probability distribution; calculating the information entropy based on the load probability distribution, and normalizing the information entropy as the fairness indicator value to characterize the degree of dispersion of workload distribution within the system.
[0070] In some embodiments of this disclosure, the aforementioned auxiliary indicator value is a specification compliance indicator value. The calculation of the auxiliary indicator value based on the first running data includes: loading a formal specification definition file, which contains specification requirements for a specific test scenario; traversing the time steps of the first running data to verify whether the state of the agent in the system under test meets the specification requirements in the formal specification definition file, and obtaining the verification result; based on the verification result, counting the total number of violation events; calculating the ratio of the total number of violations to the total running time steps or the total number of interactions of the system, and obtaining the specification compliance indicator value based on the ratio.
[0071] In this optional implementation, the first running data includes information such as the position, velocity, and actions of all agents in the system under test. The formal specification definition file is a set of formalized social norm definitions, which includes: file identifier, file description, file type, and file parameters.
[0072] In this optional implementation, the system loads formal specification definition files related to the current scenario before the test begins. At each time step of the simulation, the executing agent iterates through all agents and all loaded specifications to achieve real-time detection. For each specification, the executing agent checks whether the agent's current state satisfies its `violation_condition` (a conditional expression that violates a rule or constraint). Once satisfied, a violation event is recorded, and the total number of violations, N, is calculated based on the number of violations. violation Total system runtime N total Or the total number of interactions N timesteps The ratio NAR is shown in Equation (6).
[0073] (6)
[0074] In this optional implementation, the execution entity includes an extensible specification language interpreter capable of parsing complex specifications in formal specification definition files that contain temporal logic (such as "a response must be given within 5 seconds of a request").
[0075] In this optional implementation, a formal specification definition file containing the specific test scenario specification requirements is first loaded. Then, the first running data is traversed in time step order, and the state of the agent in the test system at each time step is checked to see if it meets the specification requirements in the file and the verification results are recorded. Next, the total number of violation events that occur in the entire trajectory is counted. Finally, the number of violations is divided by the total number of system running time steps or the total number of interactions to obtain the violation ratio. Based on this ratio, the specification compliance index value is calculated, providing a reliable implementation method for obtaining the specification compliance index value.
[0076] Optionally, the auxiliary indicator value is the standard compliance indicator value. Based on the first running data, the calculation of the auxiliary indicator value includes: pre-assigning a severity weight for each type of standard violation in the formal standard definition file; traversing the first running data, when an agent is detected to have violated the standard, calculating the single violation loss value according to the degree of violation, and calculating the weighted total violation loss in combination with the severity weight; and mapping the weighted total violation loss to a normalized value using an exponential decay function or an inverse proportional function to obtain the standard compliance indicator value.
[0077] Optionally, the auxiliary indicator value is the compliance indicator value. Based on the first operating data, the calculation of the auxiliary indicator value includes: constructing a virtual potential field based on the relative positions between the agent and environmental obstacles or other agents in the first operating data; calculating the potential field strength of the agent at each time step, where the potential field strength represents the level of violation risk; statistically analyzing the average or maximum value of the potential field strength throughout the entire operation, and calculating the compliance indicator value based on the relationship between the average or maximum value and a preset safety threshold.
[0078] In some embodiments of this disclosure, the aforementioned auxiliary index value is the adaptive recovery rate index value. The calculation of the auxiliary index value based on the first operating data includes: injecting a preset disturbance event into the test environment at a first moment during the acquisition of the first operating data; constructing a performance curve of system performance changing over time based on the first operating data; determining the performance baseline value before the disturbance, the minimum performance point value after the disturbance, and the stable performance recovery value after the disturbance based on the performance curve; and calculating the adaptive recovery rate index based on the performance baseline value, the minimum performance point value, and the stable performance recovery value to obtain the adaptive recovery rate index value.
[0079] In this optional implementation, the performance curve is time-series data of performance metrics (such as system throughput). The vertical axis represents the type and time of the disturbance event. The x-axis represents the time period. When t < t event At that time, the performance curve is at the performance baseline value P baseline The price fluctuates within a relatively stable range; when t=t event When t > t, there is a downward arrow labeled "Perturbation Injection" Z; when t > t event At that point, the performance curve plummeted to the minimum performance value P. min Then it begins to rise, eventually stabilizing at time t. end Stabilizes at performance recovery stable value P recovery Level, at Figure 2 The amount of performance degradation is indicated by falling brackets, and the amount of recovery is indicated by rising brackets.
[0080] In this optional implementation, the calculation can be performed using the first running data. The average performance within the time window is used to obtain the performance baseline value P. baseline By searching in the first running data The minimum performance value within the time window is obtained as the minimum performance point value P. min Calculated using the first running data The average performance within the time window yields the performance recovery stability value P. recovery。 The formula for calculating the adaptive recovery rate index is shown in equation (7).
[0081] (7)
[0082] In equation (7), if the performance baseline value P baseline Equal to the minimum performance value P min If the adaptive recovery rate index value is 1, then the adaptive recovery rate index value is 1.
[0083] In this optional implementation, the execution entity includes a perturbation injector, which can accurately inject various types of perturbations during test runtime according to a preset script.
[0084] In this optional implementation, before collecting the first running data, disturbance events such as node failure, communication blockage, or environmental changes are injected into the test environment at a specific time according to a preset script. Then, based on the collected first running data, a performance curve of the system performance changing over time is plotted. The performance baseline value before the disturbance, the minimum performance value after the disturbance, and the performance recovery stability value after the disturbance are extracted from the curve. The adaptive recovery rate index is calculated based on these three values to obtain the adaptive recovery rate index value, which provides a reliable implementation method for obtaining the adaptive recovery rate index value.
[0085] Optionally, the auxiliary index value is the adaptive recovery rate index value. Based on the first operating data, the calculation of the auxiliary index value includes: determining the ideal performance baseline of the multi-agent system under test in a undisturbed state; calculating the area difference between the actual performance curve reflected by the first operating data and the ideal performance baseline within a preset time window after the injection of the disturbance event; and calculating the adaptive recovery rate index value based on the ratio of the area difference to the maximum tolerable loss area, wherein the smaller the area difference, the higher the adaptive recovery rate index value.
[0086] Optionally, the auxiliary index value is the adaptive recovery rate index value. Based on the first operating data, the calculation of the auxiliary index value includes: recording the disturbance injection time of injecting a preset disturbance event into the test environment; monitoring the system performance index in the first operating data, identifying the recovery time when the system performance index recovers and stabilizes to the preset performance threshold for the first time after the disturbance; calculating the time difference between the recovery time and the disturbance injection time, and calculating the adaptive recovery rate index value based on the reciprocal or logarithmic function of the time difference.
[0087] In some embodiments of this disclosure, the above-mentioned shielding of the interaction capabilities between agents in the multi-agent system under test is achieved through at least one of the following methods: at the virtual simulation engine level, filtering out information about the position, velocity, or state of other agents in the multi-agent system under test besides the agent being tested from the observation data sent to each agent; at the virtual network level, blocking the transmission of data packets between any two agents in the multi-agent system under test; deploying each agent in the multi-agent system under test in mutually isolated independent simulation instances, and aggregating the running data of each simulation instance.
[0088] In this optional implementation, interactions between multiple agents can be simultaneously shielded through triple isolation: First, semantic filtering is performed on the observation data at the virtual simulation engine level to remove any information containing the position, velocity, or state of other agents; then, firewalls or traffic rules are used at the virtual network level to completely block the transmission of data packets between any two agents; finally, each agent is deployed to run in a completely independent simulation instance, with only the local running data of each instance being aggregated by an external scheduler, thereby achieving "zero interaction" isolation at the engine, network, and runtime layers simultaneously.
[0089] Further reference Figure 3 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a multi-agent system evaluation device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0090] like Figure 3As shown, the multi-agent system evaluation device 300 provided in this embodiment includes: a receiving unit 301, a collective value calculation unit 302, an individual value calculation unit 303, a collaborative calculation unit 304, an auxiliary calculation unit 305, and a generation unit 306. The receiving unit 301 can be configured to receive first operating data and second operating data. The first operating data is the information interaction or environmental perception sharing data between the agents in the multi-agent system under test, obtained after the multi-agent system under test has accessed a pre-configured test environment. The second operating data is the state of each agent or local environmental decision data obtained after adjusting or disabling the interaction capabilities between the agents in the multi-agent system under test in the test environment. The collective value calculation unit 302 can be configured to calculate the collective output value of the multi-agent system under test based on the first operating data. The individual value calculation unit 303 can be configured to calculate the individual output value of each agent based on the second operating data and determine the sum of the individual output values. The aforementioned collaborative calculation unit 304 can be configured to calculate a collaborative indicator value based on the sum of the collective output value and the individual output value. The aforementioned auxiliary calculation unit 305 can be configured to calculate auxiliary indicator values based on the first operational data; the auxiliary indicator values include at least one of a fairness indicator value, a norm compliance indicator value, and an adaptive recovery rate indicator value. The aforementioned generation unit 306 can be configured to generate multi-dimensional quantitative indicator values for the multi-agent system under test based on the collaborative indicator value and the auxiliary indicator values.
[0091] In some embodiments of this disclosure, the multi-agent system evaluation device 300 further includes an authentication unit (not shown in the figure), which is configured to generate a visual authentication report based on the multi-dimensional quantitative index values.
[0092] In some embodiments of this disclosure, the collaborative calculation unit 304 is configured to: subtract the sum of individual output values from the collective output value to obtain a net gain value; determine a collaborative index value based on the ratio of the net gain value to the sum of individual output values; wherein, if the sum of individual output values is zero and the collective output value is greater than zero, the collaborative index value is set to a preset saturation value.
[0093] In some embodiments of this disclosure, the above-mentioned auxiliary index value is a fairness index value, and the above-mentioned auxiliary calculation unit 305 is configured to: extract the resource request events and corresponding waiting times of each agent for shared resources from the first running data; calculate the statistical coefficient reflecting the dispersion of resource allocation based on the waiting time distribution of all agents, and use the statistical coefficient as a fairness index value.
[0094] In some embodiments of this disclosure, the aforementioned auxiliary index value is a specification compliance index value, and the aforementioned auxiliary calculation unit 305 is further configured to: load a formal specification definition file, which contains specification requirements for a specific test scenario; traverse the time steps of the first running data to verify whether the state of the agent in the system under test meets the specification requirements in the formal specification definition file, and obtain the verification result; based on the verification result, count the total number of violation events; calculate the ratio of the total number to the total running time steps or the total number of interactions of the system, and obtain the specification compliance index value based on the ratio.
[0095] In some embodiments of this disclosure, the aforementioned auxiliary index value is the adaptive recovery rate index value, and the aforementioned auxiliary calculation unit 305 is further configured to: inject a preset disturbance event into the test environment at a first moment during the process of acquiring the first operating data; construct a performance curve of system performance changing over time based on the first operating data; determine the performance baseline value before the disturbance occurs, the minimum performance point value after the disturbance occurs, and the stable performance recovery value after the disturbance ends based on the performance curve; calculate the adaptive recovery rate index based on the performance baseline value, the minimum performance point value, and the stable performance recovery value to obtain the adaptive recovery rate index value.
[0096] like Figure 4 As shown, the multi-agent system evaluation system 400 provided in this embodiment includes: an access module 401, an orchestration module 402, a test control module 403, a core computing engine 404, and a report generation module 405. The access module 401 is configured to receive the multi-agent system under test (MAS) packaged as a standardized deployment unit. The orchestration module 402 is configured to dynamically create an isolated test environment and deploy the MAS and a standardized virtual simulation platform within the test environment. The test control module 403 is configured to control the virtual simulation platform to execute a test process, which includes: enabling inter-agent interaction in the MAS to perform collaborative operation testing and collecting first operation data in a first time period; and disabling inter-agent interaction in the MAS to perform independent operation testing and collecting second operation data in a second time period. The core computing engine 404 is configured to receive the first and second operation data, and based on the first and second operation data, execute the multi-agent system evaluation method described above to calculate multi-dimensional quantitative indicators, including collaborative efficiency. The report generation module 405 is configured to generate a visual certification report based on multi-dimensional quantitative indicator values.
[0097] In this embodiment, the standardized deployment unit can be a unit implemented by deployment methods such as virtual machines, direct executable files, and container images.
[0098] In this embodiment, the multi-agent system evaluation system 400 is a plug-and-play virtual simulation testing platform that can serve as a standardized examination room. In one example, the multi-agent system evaluation system 400 is a system based on containerization technology and a client-server architecture. The server runs the physical simulation environment and the core data acquisition and monitoring modules; the client (the MAS under test) packages the MAS under test into a standard Docker image. This image needs to expose an endpoint that conforms to a defined gRPC or RESTful API.
[0099] The Multi-Agent System Evaluation System 400 follows the API specification (example): GetObservations(agent_ids): The server calls this API to send sensor observation data of the specified agent to the MAS; SetActions(actions): The MAS calls this API to submit the next action instructions for all its agents to the server. This clear API definition allows any third-party MAS to be easily integrated.
[0100] The virtual simulation platform's built-in test scenarios include: "Digital Rwanda," a scenario containing limited food resources that slowly regenerate over time. Agents need energy to survive, which can be replenished by foraging. The key parameter in this scenario is the ratio of resource regeneration rate to the agent's energy consumption rate; adjusting this parameter controls survival pressure. "Tower of Babel" is a complex assembly task requiring different types of parts. The scenario includes two agents: Agent A can only see and pick up type A parts, and Agent B can only see and pick up type B parts. However, the final assembly station requires both types of parts (A and B). They do not have a direct communication channel; they must infer each other's intentions and states by observing each other's behavior and location (e.g., seeing the other linger near a parts bin) to achieve tacit cooperation.
[0101] The workflow of the Multi-Agent System Evaluation System 400 is as follows: After the user completes the upload and configuration through the front-end interface, the workflow of the Multi-Agent System Evaluation System 400 starts: The access module 401 dynamically creates an isolated test namespace based on Kubernetes or similar technologies; in this space, it pulls and starts the server-side container of the test suite; the orchestration module 402 pulls and starts the MAS client container uploaded by the user; the test control module 403 starts the test script through the orchestrator, and the monitoring module collects performance data and trajectory logs in real time; after the test, an analysis job is triggered, which runs the core computing engine 404; the core computing engine 404 processes the log data and calculates the C-IQ (Collective Intelligence Quotient) scores. Then, using a report template, the results are rendered into an authentication report in PDF or HTML format and sent to the report generation module 405; the report generation module 405 notifies the user via email or webhook and automatically destroys the test environment and reclaims resources.
[0102] In this embodiment, the generated certification report is not just a score, but also a diagnostic report. Its content includes: a total C-IQ score and its sub-scores radar chart, performance curves, key event replay information, and diagnostic recommendations. The performance curves are detailed curves showing how various indicators (such as throughput and conflict rate) change over time. The key event replay information includes automatically captured and visualized "highlights" (such as a successful complex collaboration) and "darkest moments" (such as a systemic crash) during the testing process. The diagnostic recommendations are automatically generated by an expert system or LLM based on the score weaknesses, providing targeted improvement suggestions. For example: If your system scores low on GWT, it is recommended to optimize the resource queuing algorithm.
[0103] The following example, using the Tower of Babel scenario, illustrates the workflow of the multi-agent system evaluation system implemented in this paper:
[0104] [Test Scenario] In a 100m x 100m virtual warehouse, there are two parts bins (bin A contains only red cubes, and bin B contains only blue spheres) and a central assembly table. The task is to match a red cube with a blue sphere and move them to the assembly table to complete one scoring cycle.
[0105] The MAS under test consists of four robotic agents. Agent_1 and Agent_2 are set as type A (can only sense and pick up red squares), and Agent_3 and Agent_4 are set as type B (can only sense and pick up blue spheres). There is no direct communication channel configured between them.
[0106] Step 1: User Submission and System Deployment
[0107] The user packages their developed MAS (Multi-Agent System) for testing into a Docker image conforming to this public API specification and uploads it through the authentication system's web frontend. This image exposes a gRPC endpoint at mas-service:50051. Upon receiving the request, the Multi-Agent System Evaluation System 400 creates a new namespace, test-session-xyz, in the Kubernetes cluster and deploys two core services: isaac-sim-server and mas-under-test. isaac-sim-server is used to run the "Tower of Babel" simulation environment, and mas-under-test is used to run the user-uploaded MAS image for testing.
[0108] Step 2: Test Execution and Data Acquisition
[0109] The total test duration is set to T_total = 3600 simulation time steps. The main control script starts and begins the test loop; after the test ends, the orchestrator triggers the C-IQ total score calculation job, which loads the trajectory.csv file and sequentially calls the four calculation modules for collaboration index, fairness index, standard compliance index, and adaptive recovery rate index.
[0110] (a) Calculation of collaboration index values:
[0111] = Number of [red cube + blue sphere] pairs successfully delivered to the assembly table
[0112] According to the log statistics, O collective There are 52 pairs. The system automatically reruns the test in independent mode (disabling the visibility of other agents). Since no single agent can complete pairing independently, its independent output is O. individual,i All are 0. Therefore, O sum = 0, collaboration metric value SE = 1.0, indicating that the task has strong collaboration dependencies.
[0113] (b) Calculation of fairness index values:
[0114] The assembly station is the only mutually exclusive shared resource. An event detector scans the trajectory, identifying "start waiting" events where each agent arrives near the assembly station (within a 2-meter radius) but cannot place an item because the station is occupied, and "end waiting" events where an item is successfully placed. Agent_1's total waiting time W1 = 35 steps; Agent_2's total waiting time W2 = 41 steps; Agent_3's total waiting time W3 = 38 steps; Agent_4's total waiting time W4 = 32 steps. Based on the formula for calculating the fairness metric, the fairness metric value GWT is calculated to be 0.058. This is a very low value, indicating that resource allocation is very fair.
[0115] (c) Calculation of the National Accuracy Ratio (NAR) value:
[0116] Load the specification file for the Tower of Babel scenario, which contains two key specifications: NO_COLLISION: The distance between any two agents must be greater than 0.5 meters; ITEM_PROTECTION: When an agent is holding an item, its movement speed must not exceed 1.0 m / s.
[0117] The scan revealed three minor collisions (distance less than 0.5 meters) across the entire 3600 time steps. No events of speeding while carrying an item were detected. violation = 3, then the compliance index value NAR = 1 - (3 / 3600) = 0.9992, which is an extremely high score, indicating that the system behavior is very predictable and safe.
[0118] (d) Calculation of Adaptive Recovery Rate (ARR):
[0119] This refers to the number of items successfully paired within the past 100 time steps. (See reference...) Figure 2 Within the first [1700, 1800] steps of the disturbance, the average performance P baseline =1.5 pairs / 100 steps. At t=1800, Agent_1 is forcibly offline. Due to the reduction of type A agents by half, the system output drops sharply. Within the range of [1800, 1900] steps, the lowest performance point P is reached. min =0.6 pairs / 100 steps. After a period of adaptation, Agent_2 (Type A) began to take on more work and formed a new cooperative pattern with Agent_B. Within the last 3500, 3600 steps of the test, the average performance recovered to P. recovery =1.1 pairs / 100 steps. Adaptive recovery rate index value ARR = (1.1 - 0.6) / (1.5 - 0.6) = 0.5 / 0.9 ≈ 0.556`.
[0120] Step 3: C-IQ Total Score Calculation and Certification Report Generation
[0121] Assuming that for collaborative, intensive tasks like the Tower of Babel, the weights are set as {w1_SE=0.4, w2_GWT=0.2, w3_NAR=0.1, w4_ARR=0.3}. C-IQ = 0.4 * 1.0 + 0.2 * (1 - 0.058) + 0.1 * 0.9992 + 0.3 * 0.556 = 0.4 + 0.1884 + 0.09992 + 0.1668 = 0.85512. The final C-IQ score (out of 100) is 85.5.
[0122] [Generate Certification Report] The system calls report generation module 405 to populate all the above data and charts into an HTML template, generating a final report with both text and graphics. The report summary will give the following conclusion: "Certification Conclusion: The tested MAS demonstrates excellent collective intelligence in the Tower of Babel scenario, with a C-IQ score of 85.5. Its advantages lie in its extremely high collaboration efficiency and fairness. The main improvement point of the system lies in its adaptive recovery capability when facing core node failures, with an ARR score of 55.6, indicating that its collaboration mode has a certain dependence on specific individuals. It is recommended to optimize its task dynamic reallocation strategy to enhance system resilience."
[0123] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0124] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0125] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their patterns are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0126] like Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0127] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0128] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as multi-agent system evaluation methods. For example, in some embodiments, the multi-agent system evaluation method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the multi-agent system evaluation method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform multi-agent system evaluation methods by any other suitable means (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to the processor or controller of a general-purpose computer, special-purpose computer, or other programmable multi-agent system evaluation apparatus, such that when executed by the processor or controller, the program code causes the patterns / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0135] The foregoing description of specific exemplary embodiments of this disclosure is for illustrative and explanatory purposes. These descriptions are not intended to limit this disclosure to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of this disclosure and their practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of this disclosure, as well as various different choices and variations. The scope of this disclosure is intended to be defined by the claims and their equivalents.
Claims
1. A method for evaluating a multi-agent system, the method comprising: The system receives first running data and second running data. The first running data is the information interaction or environmental perception sharing data between the agents in the multi-agent system under test obtained after the system under test accesses the pre-configured test environment. The second running data is the state of each agent or local environmental decision data of each agent obtained after adjusting or disabling the interaction capabilities between the agents in the multi-agent system under test in the test environment. Based on the first operational data, calculate the collective output value of the multi-agent system under test; Based on the second operational data, calculate the individual output value of each agent and determine the sum of the individual output values; The collaboration index value of the collaboration efficiency index is calculated based on the sum of the collective output value and the individual output value. Based on the first operational data, auxiliary indicator values are calculated; the auxiliary indicator values include at least one of the following: fairness indicator value, standard compliance indicator value, and adaptive recovery rate indicator value. Based on the collaboration index value and the auxiliary index value, multi-dimensional quantitative index values of the multi-agent system under test are generated.
2. The method according to claim 1, further comprising: Based on the multi-dimensional indicator values, a visual authentication report is generated.
3. The method according to claim 1 or 2, wherein, The calculation of the collaboration efficiency index value based on the sum of the collective output value and the individual output value includes: Subtracting the sum of the individual outputs from the collective output value yields the net gain value. The collaboration index value is determined based on the ratio of the net gain value to the sum of the individual output values. If the sum of the individual output values is zero and the collective output value is greater than zero, then the collaboration index value is set to a preset saturation value.
4. The method according to claim 1 or 2, wherein, The auxiliary indicator value is a fairness indicator value, and the calculation of the auxiliary indicator value based on the first operating data includes: From the first running data, extract the resource request events for shared resources of each agent and the corresponding waiting time; Based on the waiting time distribution of all agents, a statistical coefficient reflecting the dispersion of resource allocation is calculated, and the statistical coefficient is used as the fairness index value.
5. The method according to claim 1, wherein, The auxiliary indicator value is the standard compliance indicator value, and the calculation of the auxiliary indicator value based on the first operational data includes: Load the formal specification definition file, which contains specification requirements for a specific test scenario; Traverse the time steps of the first running data to verify whether the state of the agent in the system under test meets the specification requirements in the formal specification definition file, and obtain the verification result; Based on the verification results, the total number of violations will be counted. Calculate the ratio of the total number of times to the total system runtime steps or total number of interactions, and obtain the specification compliance index value based on this ratio.
6. The method according to claim 1, characterized in that, The auxiliary indicator value is the adaptive recovery rate indicator value, and the calculation of the auxiliary indicator value based on the first operating data includes: During the process of acquiring the first running data, a preset disturbance event is injected into the test environment at the first moment; Based on the first operational data, construct a performance curve of the system performance over time; Based on the performance curve, determine the performance baseline value before the disturbance, the minimum performance value after the disturbance, and the performance recovery stability value after the disturbance ends. Based on the performance baseline value, the performance minimum value, and the performance recovery stability value, the adaptive recovery rate index is calculated to obtain the adaptive recovery rate index value.
7. A multi-agent system evaluation device, the device comprising: The receiving unit is configured to receive first operating data and second operating data. The first operating data is information interaction or environmental perception sharing data between agents in the multi-agent system under test obtained after the multi-agent system under test is connected to a pre-configured test environment. The second operating data is the state of each agent or local environmental decision data of each agent obtained after adjusting or blocking the interaction capabilities between agents in the multi-agent system under test in the test environment. The collective value calculation unit is configured to calculate the collective output value of the multi-agent system under test based on the first running data. The individual value calculation unit is configured to calculate the individual output value of each agent based on the second running data, and determine the sum of the individual output values; The collaborative computing unit is configured to calculate a collaborative index value for the collaborative efficiency index based on the sum of the collective output value and the individual output value. An auxiliary calculation unit is configured to calculate auxiliary indicator values based on the first operating data; the auxiliary indicator values include at least one of a fairness indicator value, a standard compliance indicator value, and an adaptive recovery rate indicator value. The generation unit is configured to generate multi-dimensional quantitative index values for the multi-agent system under test based on the collaborative index values and the auxiliary index values.
8. A multi-agent system evaluation system, characterized in that, include: The access module is configured to receive a multi-agent system under test encapsulated as a standardized deployment unit. The orchestration module is configured to dynamically create isolated test environments and deploy the multi-agent system under test and a standardized virtual simulation platform within those test environments. The test control module is configured to control the virtual simulation platform to execute a test process, which includes: enabling inter-agent interaction in the multi-agent system under test in a first time period to perform collaborative operation test and collect first operation data; and disabling inter-agent interaction in the multi-agent system under test in a second time period to perform independent operation test and collect second operation data. The core computing engine is configured to receive the first running data and the second running data, and based on the first running data and the second running data, execute the method as described in any one of claims 1 to 6 to calculate multi-dimensional quantitative indicators, including collaboration efficiency. The report generation module is configured to generate a visual authentication report based on the multi-dimensional quantitative indicator values.
9. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.