Performance evaluation method of intelligent agent and electronic device

By constructing a multi-dimensional performance evaluation framework, acquiring communication, decision-making, and state data of agents, and conducting real-time evaluation and cross-validation, the problem of single evaluation dimensions and reliance on manual anomaly diagnosis in multi-agent systems is solved, enabling accurate root cause localization and efficient diagnosis of the performance of multi-agent systems.

CN122432007APending Publication Date: 2026-07-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-06-18
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing performance evaluations of multi-agent systems lack quantitative assessment of the value of communication content, collaborative stability assessments focus on output results and lack fine-grained behavioral analysis, making it impossible to monitor fairness during the process, and anomaly diagnosis relies on manual analysis with ambiguous root causes.

Method used

By acquiring the communication data, decision data, and state data of the intelligent agent, a multi-dimensional performance evaluation framework is constructed, including a communication effectiveness evaluation module, a collaborative decision stability module, and a collaborative fairness state monitoring module. The framework is evaluated in real time and cross-validated to generate a performance evaluation report.

Benefits of technology

It enables precise root cause localization of the performance of multi-agent systems, improves the comprehensiveness of the assessment and the accuracy of the diagnosis, and solves the problems of single assessment dimensions and reliance on manual anomaly diagnosis in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432007A_ABST
    Figure CN122432007A_ABST
Patent Text Reader

Abstract

The application discloses a kind of performance evaluation method and electronic equipment of intelligent agent, it is related to artificial intelligence technical field, including obtaining the communication data, decision data and state data of intelligent agent, respectively obtain the communication effectiveness evaluation result, collaborative decision stability evaluation result and cooperation fair state evaluation result of intelligent agent;In response to any one of the above three evaluation results is abnormal result, then using communication data, decision data and state data cross-validation obtain the abnormal reason of intelligent agent, and according to evaluation result and abnormal reason obtain the performance evaluation report of intelligent agent, solve the problem that related technology multi-agent system performance evaluation dimension is single, abnormal diagnosis relies on artificial and root cause is fuzzy, by fusing communication effectiveness, collaborative decision stability and cooperation fairness three-dimensional dynamic evaluation and cross-validation, realize the accurate root cause positioning of multi-agent system performance anomaly, improve the comprehensiveness of evaluation and the accuracy of diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for evaluating the performance of intelligent agents and an electronic device. Background Technology

[0002] With the widespread application of multi-agent systems in fields such as autonomous driving, intelligent scheduling, and distributed robotics, the evaluation of their collaborative performance has become crucial for system optimization and maintenance.

[0003] In related technologies, the MARLlib framework is usually used for multi-dimensional evaluation, but the following problems still exist: (1) lack of quantitative evaluation of the value of communication content; (2) the evaluation of collaborative stability focuses on the output of results and lacks fine-grained behavior analysis; (3) the fairness evaluation is mostly post-event accounting and cannot achieve in-process monitoring; (4) the anomaly location depends on manual analysis and lacks an automated root cause location mechanism, which urgently needs to be solved. Summary of the Invention

[0004] This invention provides a performance evaluation method for intelligent agents, which at least solves the problems of single performance evaluation dimension, reliance on manual diagnosis, and unclear root causes in the performance evaluation of multi-agent systems in related technologies.

[0005] This invention provides a method for evaluating the performance of an intelligent agent, comprising: acquiring communication data, decision data, and state data of the intelligent agent; obtaining communication effectiveness evaluation results, collaborative decision stability evaluation results, and collaborative fairness state evaluation results of the intelligent agent based on the communication data, the decision data, and the state data; responding to any one of the evaluation results being an anomalous, performing cross-validation using the communication data, the decision data, and the state data to obtain a cross-validation result, determining the cause of the anomalousness based on the cross-validation result, and obtaining a performance evaluation report of the intelligent agent based on the evaluation results and the cause of the anomalousness.

[0006] The present invention also provides a performance evaluation device for an intelligent agent, comprising: The first acquisition module is used to acquire the agent's communication data, decision data, and state data; The second acquisition module is used to obtain the communication effectiveness evaluation result, collaborative decision stability evaluation result, and collaborative fairness state evaluation result of the agent based on the communication data, the decision data, and the state data. The third acquisition module is used to respond to any one of the communication effectiveness assessment result, the collaborative decision stability assessment result, and the collaborative fairness state assessment result being an abnormal result, by performing cross-validation using the communication data, the decision data, and the state data to obtain a cross-validation result, obtaining the cause of the abnormality of the agent based on the cross-validation result, and obtaining a performance assessment report of the agent based on the assessment result and the cause of the abnormality.

[0007] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described intelligent agent performance evaluation methods.

[0008] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described intelligent agent performance evaluation methods.

[0009] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described intelligent agent performance evaluation methods.

[0010] This invention obtains communication effectiveness evaluation results, collaborative decision-making stability evaluation results, and collaborative fairness state evaluation results for intelligent agents by acquiring their communication data, decision-making data, and state data. In response to any of these three evaluation results being an anomaly, cross-validation is performed using the communication data, decision-making data, and state data to determine the cause of the anomaly. A performance evaluation report for the intelligent agent is then generated based on the evaluation results and the cause of the anomaly. This invention solves the problems of single-dimensional performance evaluation of multi-agent systems, reliance on manual diagnosis, and ambiguous root causes in related technologies. By integrating dynamic evaluation and cross-validation of the three dimensions of communication effectiveness, collaborative decision-making stability, and collaborative fairness, it achieves precise root cause localization of performance anomalies in multi-agent systems, improving the comprehensiveness of the evaluation and the accuracy of the diagnosis. Attached Figure Description

[0011] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a performance evaluation method for an intelligent agent provided in an embodiment of the present invention; Figure 2This is an overall flowchart of multi-agent collaboration multi-dimensional performance evaluation provided in one embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the stability assessment of collaborative decision-making according to an embodiment of the present invention. Figure 4 An exception feedback flowchart is provided for one embodiment of the present invention; Figure 5 This is a block diagram of a performance evaluation device for an intelligent agent according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0014] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0015] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0016] Specifically, before introducing the embodiments of the present invention, we will first introduce the background technology of the present invention and the technical defects existing in the related technologies.

[0017] With the rapid development of artificial intelligence technology, intelligent agents have been widely used in many fields such as autonomous driving, financial analysis, and medical assistance. Their complexity and the diversity of application scenarios continue to increase, and scientifically and systematically evaluating the value of intelligent agents has become a key issue that urgently needs to be addressed.

[0018] Intelligent agent value assessment is a process of quantitative and qualitative analysis of an intelligent agent's capabilities, efficiency, reliability, and social contribution in a task and environment through specific indicators and methods. It needs to take into account both technical performance and non-technical dimensions such as ethics and economics. Among them, the technical aspect focuses on the efficiency, quality, environmental adaptability, and generalization ability of functional implementation; the non-technical aspect covers ethical compliance, cost-effectiveness, and decision interpretability.

[0019] Current assessments are based on a combination of quantitative and qualitative methods. Quantitative methods rely on quantifiable indicators such as accuracy and response time, while qualitative methods focus on complex dimensions such as user experience and social impact. This is supplemented by third-party evaluations and expert reviews to build a multi-perspective system. At the same time, the assessment system needs to evolve dynamically with technological iterations and social needs, possessing flexibility and foresight. Its scientific construction is a systematic project that supports the healthy development of intelligent agents.

[0020] For example, in the field of multi-agent reinforcement learning, MARLlib, as an open-source framework, has built a systematic evaluation system. Its core methods cover multi-dimensional indicators and multi-level testing strategies. For agent collaboration effectiveness, core indicators such as task completion and success rate are used to quantify goal achievement capabilities, while efficiency indicators (such as interaction steps and resource consumption) are combined to evaluate decision-making economy. Robustness testing verifies the system's fault tolerance capabilities under abnormal environments such as noise interference and information loss, and indicators such as context memory length and planning reasoning ability are used to measure the consistency between complex task decomposition and long-term goals. For large-model-driven multi-agent systems, MARLlib further introduces a dual-dimensional evaluation framework of autonomy and consistency. Through interactive analysis of core functional modules such as goal decomposition, task orchestration, and multi-agent collaboration, it quantifies the dynamic adaptability of the system architecture and the consistency of results, providing a reproducible technical path for evaluating agent collaboration effectiveness in complex scenarios. This platform aims to overcome the common problems in MARL evaluation, such as a single benchmark and difficulty in reproducing results, providing a more comprehensive and reliable measure of algorithm performance.

[0021] Therefore, the aforementioned multi-agent evaluation method based on MARLlib lacks a standard for measuring the value of communication content, making it difficult to quantify the actual contribution of each message to the final decision, and thus failing to effectively distinguish between redundant or invalid communication. Regarding collaborative stability, it leans more towards measuring the stability of the output results, failing to perform real-time, fine-grained analysis based on the collaborative behavior of multiple agents, and lacks a multi-dimensional in-process fairness monitoring mechanism. When analyzing the causes of anomalies, it relies more on manual analysis of logs and indicator curves to troubleshoot problems, rather than establishing an automated, cross-module mechanism for locating the root causes of anomalies.

[0022] Therefore, based on the aforementioned technical problems, this invention obtains abnormal feedback by jointly monitoring communication effectiveness, collaborative decision-making stability, and collaborative fairness, and automatically reverses the source of the abnormal behavior when abnormal feedback is received, and provides relevant handling suggestions.

[0023] The embodiments of the present invention provide a performance evaluation method for an intelligent agent, and the method is described in detail in conjunction with the execution flow of the performance evaluation method for an intelligent agent.

[0024] Specifically, Figure 1 This is a flowchart illustrating a performance evaluation method for an intelligent agent provided in an embodiment of the present invention. like Figure 1 As shown, the performance evaluation method for intelligent agents includes the following steps: In step S101, the agent's communication data, decision data, and state data are acquired.

[0025] According to an embodiment of the present invention, acquiring communication data, decision data, and status data of an intelligent agent includes: determining attribute information to be collected, wherein the attribute information to be collected includes at least one of message identifier, sending intelligent agent, receiving intelligent agent, message type, message content characteristics, and message sending time; collecting data from the attribute information to be collected to obtain communication data, and querying the data of the receiving intelligent agent using a preset decision status query interface to obtain decision data; and acquiring task execution logs and task monitoring logs generated during the operation of the intelligent agent to obtain status data based on the task execution logs and task monitoring logs.

[0026] The preset decision status query interface can be selected by those skilled in the art based on actual decision-making performance and needs, and is not specifically limited here.

[0027] Specifically, this invention constructs a closed-loop, multi-dimensional performance evaluation framework for multi-agent cooperative systems, such as... Figure 2 As shown, the system consists of three core evaluation modules and a unified anomaly feedback module. The three core evaluation modules include a communication effectiveness evaluation module, a collaborative decision-making stability module, and a collaborative fairness monitoring module. These modules collect raw behavioral data during system operation in parallel, conduct independent special evaluations, and report the evaluation results to the anomaly feedback module in real time. As the decision-making center, the anomaly feedback module, after receiving the uploaded evaluation results, performs cross-validation of data across modules to locate the root cause of system anomalies and develop optimization strategies. Finally, it outputs a structured evaluation report, thus forming a complete closed loop of "perception, analysis, diagnosis, and feedback".

[0028] The communication effectiveness assessment module calculates communication efficiency-related indicators during the communication process through "communication-decision correlation" and issues an early warning when the effective communication rate is low. The collaborative decision stability module comprehensively evaluates the decision stability of the agents during the collaborative process from two dimensions: synchronization (group goal consensus) and stability (consensus stability) and issues an early warning when the decision is unstable. The collaborative fairness monitoring module moves from "post-event accounting" to "in-event dynamic monitoring," enabling real-time monitoring of the agents' gains and resources after each round of tasks, outputting the deviation of gains and resources respectively, and detecting abnormal issues such as long-term low gains, fairness instability, and fairness deterioration of each agent through indicators such as the magnitude of deviation changes, and issuing early warnings. The anomaly feedback module integrates the above three modules to perform data cross-validation, locks the root cause of anomalies according to priority, generates optimization strategies based on the root cause of anomalies, and outputs evaluation reports, which greatly improves the efficiency of anomaly handling.

[0029] Specifically, this invention first acquires the communication data, decision data, and state data of the agents, thereby providing structured, time-aligned, and semantically clear raw behavior logs for the subsequent three evaluation modules: communication effectiveness, collaborative decision stability, and collaborative fairness. Communication data refers to the message records exchanged between agents during collaboration, used to analyze the effectiveness and efficiency of information transmission. As shown in Table 1, communication data mainly includes at least one of the following: message identifier, sending agent, receiving agent, message type, message content characteristics, and message sending time. Decision data reflects the behavioral choices or strategy outputs made by each agent at a specific moment and is the core basis for evaluating collaborative quality. It can be queried through a preset decision state query interface. The system collects data from intelligent agents; state data characterizes the agent's gains and resource consumption during task execution, serving as a direct basis for fairness monitoring. This data can be obtained from task execution logs and task monitoring logs generated during agent operation. Secondly, to ensure that the above three types of data can be jointly analyzed, the system also needs to perform the following preprocessing, such as time synchronization, format standardization, noise reduction, and filtering. Thus, this invention, through an automated and structured log collection mechanism, can comprehensively, in real-time, and with fine granularity capture key behavioral evidence in the collaboration process from three dimensions: information interaction (communication), behavioral output (decision-making), and input-output (state). This provides a solid and reliable data foundation for subsequent multi-dimensional performance evaluation and anomaly diagnosis.

[0030] Table 1

[0031] Thus, by collecting three key behavioral data—communication, decision-making, and state—in a structured manner, a comprehensive, accurate, and low-intrusive perception of the operation process of a multi-agent system was achieved, providing a high-fidelity, time-aligned data foundation for subsequent multi-dimensional performance evaluation and anomaly root cause localization.

[0032] In step S102, the communication effectiveness evaluation result, collaborative decision stability evaluation result, and collaborative fairness state evaluation result of the agent are obtained based on the communication data, decision data, and state data.

[0033] According to one embodiment of the present invention, obtaining communication effectiveness evaluation results, collaborative decision stability evaluation results, and collaborative fairness state evaluation results of an intelligent agent based on communication data, decision data, and state data includes: constructing a correlation time window for communication data, and analyzing whether the communication data triggers corresponding decision data within the time window, so as to obtain communication effectiveness evaluation results based on the decision results; constructing a decision time sequence based on the decision data, calculating the decision synchronization degree and decision stability of the intelligent agent based on the decision time sequence, and calculating the collaborative stability comprehensive score of the intelligent agent using a weighted summation method based on the decision synchronization degree and decision stability, so as to obtain collaborative decision stability evaluation results based on the collaborative stability comprehensive score; obtaining the current round's profit deviation degree, current round's resource deviation degree, previous round's profit deviation degree, and previous round's resource deviation degree based on the state data, and obtaining the profit change magnitude based on the current round's profit deviation degree and previous round's profit deviation degree, and obtaining the resource change magnitude based on the current round's resource deviation degree and previous round's resource deviation degree, and obtaining the collaborative fairness state evaluation results based on the profit change magnitude and resource change magnitude.

[0034] Specifically, this invention constructs a multi-dimensional, process-oriented, and quantifiable multi-agent performance evaluation system by integrating communication data, decision data, and state data. It deeply characterizes the operational quality of the agent system from three key dimensions: information effectiveness, group synergy, and collaborative fairness.

[0035] Specifically, in existing multi-agent systems, communication effectiveness assessment typically only counts the number of communications or bandwidth usage, without determining whether the communication actually participates in the decision-making process. This results in high communication load but low collaboration efficiency in the multi-agent system. To address this, this invention provides a "communication-decision association" communication effectiveness assessment module that calculates communication efficiency-related indicators during the communication process and issues an early warning when the effective communication rate is low.

[0036] This invention first collects key attribute information related to each communication data of the intelligent agent, mainly including message identifier, sending intelligent agent, receiving intelligent agent, message type, message content characteristics, and message sending time, as shown in Table 1 above. The message identifier is the message ID, the sending intelligent agent is the sender identifier, the receiving intelligent agent is the receiver identifier, the message sending time is the timestamp, the message content characteristics are key fields, category identifiers or summaries, and the message type is state synchronization / intent notification / request / instruction.

[0037] Secondly, for each receiving agent, relevant decision information is obtained by actively calling the preset decision status query interface provided by the agent. The collection frequency is determined by a combination of event triggering and periodic polling. When the preset decision status query interface supports "decision change notification," that is, once the agent makes a decision, it will notify the agent through the preset decision status query interface, prioritizing the use of the interface event triggering method to obtain the decision status. When the preset decision status query interface does not support change notification, a periodic collection method is used as a supplement, with the collection frequency based on the agent's historical average decision cycle. Configure settings to ensure that changes in decision-making can be captured in a timely manner.

[0038] It should be noted that if frequent changes in decision-making are detected during the data collection process, the data collection frequency can be appropriately increased to ensure that no changes in decision-making are missed; if the decision-making remains stable over a long period, the data collection frequency should be appropriately reduced to avoid wasting resources and affecting the normal operation of the agent. The key attributes that need to be collected are shown in Table 2. Table 2

[0039] Secondly, an associated time window is established for each piece of communication data, and the communication data is analyzed within the time window to determine whether it triggers corresponding decision data. In other words, it is observed whether the relevant receiving agent undergoes a decision change related to the communication data within the time window. The time window value W is defined in this invention as:

[0040] in, This is the lower limit of the time window. P90 is the upper limit of the time window to prevent it from being too large or too small. P90 is the 90th percentile delay value (obtained by statistically analyzing the delay distribution from receiving the message to the first relevant decision change in historical messages for each receiving agent). k is the amplification factor used to expand the time window to cover the tail of the delay. It is generally in the range of 1-2.5. Increasing k will widen the time window and associate unrelated decisions with messages, increasing the burden of computation, storage and deduplication. However, if k is too small, it is easier to miss real triggers with large delays. Therefore, it needs to be dynamically adjusted according to the actual business situation.

[0041] Therefore, by constructing causal relationships between communication and decision-making, quantifying the stability of group decision-making in two dimensions, and monitoring dynamic fairness based on deviation changes, a refined, process-oriented, and interpretable comprehensive performance evaluation of multi-agent systems in three key dimensions—communication effectiveness, collaborative stability, and collaborative fairness—is achieved.

[0042] According to one embodiment of the present invention, analyzing whether communication data triggers corresponding decision data within a time window to obtain a communication effectiveness evaluation result based on the decision result includes: if the communication data triggers corresponding decision data respectively, associating the message content in the agent's communication data with the corresponding decision data based on a preset matching strategy to obtain data tags for the communication data respectively; and obtaining a communication effectiveness evaluation result based on the data tags.

[0043] According to one embodiment of the present invention, the message content in the communication data of an intelligent agent is associated with the corresponding decision data based on a preset matching strategy to obtain data tags for the communication data, including: determining whether there is structured data in the message content and the decision data; if there is structured data in the message content and the decision data, extracting a first target field of the message content and a second target field of the decision data, and detecting the matching degree between the first target field and the second target field, and obtaining data tags for multiple communication data based on the matching degree; otherwise, constructing an association between the message content and the decision data based on a preset matching method; identifying preset keywords in the message content and decision information corresponding to the preset keywords in the decision data; and obtaining data tags for the communication data based on the preset keywords and the decision information corresponding to the preset keywords.

[0044] The preset matching strategy can be selected by those skilled in the art according to actual needs, and no specific limitations are made here.

[0045] Specifically, if the receiving agent makes a decision related to the communication data within the time window, the decision record entry of the receiving agent is retrieved within the window period [recv_time, recv_time+W] of each communication data, where recv_time is the data reception time and W is the window value defined above. After collecting relevant information, the present invention adopts a preset matching strategy, namely a multi-level matching mechanism, to realize the association between message content and agent decision data, and assigns relevant data tags to each communication data. The tags of each communication data can be preset as: effective communication (explicitly triggers decision), potentially effective communication (may trigger decision), and invalid communication (does not trigger decision).

[0046] Specifically, the system pre-builds and maintains a "message type - decision type" mapping table. The mapping relationship can be determined by manual configuration or by fixing it during the design phase. The message type is the type identifier of the input message, and the decision type is the decision type identifier preset by the agent, such as state update message - matching state adjustment decision.

[0047] Furthermore, the system first parses the content of the communication message and the decision data within the corresponding time window to determine whether it contains parsable structured fields. If structured data exists in the message content and decision data, the system extracts the first target field in the message content, namely the core business field (denoted as message.content_keys), and the second target field in the agent's decision data, namely the related business field (denoted as decision.related_fields), thereby performing a precise field-level comparison.

[0048] Secondly, the matching degree between the first target field and the second target field is detected. Based on the matching degree, data tags for multiple communication data are obtained. The matching rules are as follows: if the preset field matching conditions such as message.content_keys.task_id==decision.related_task_id and message.content_keys.action==decision.decision_value.action are met, the decision is determined to be a strong match with the message type, and the triggering status is marked as "explicitly triggered decision", that is, the data tag is valid communication; if none of them match, the status is marked as "not triggered decision", that is, the data tag is "invalid communication"; if only a partial match is found, the status is marked as "potentially triggered decision", that is, the data tag is "potentially valid communication".

[0049] Optionally, for unstructured messages without strictly defined structured fields, a pre-defined matching method is used, namely string inclusion matching or keyword matching, to establish the association between message content and decision data. The matching rule is as follows: if the text content of the message contains a pre-defined keyword (such as "move"), and the decision content generated by the agent is a decision corresponding to the keyword (such as moving to (...)), then the triggering state is marked as "explicitly triggering a decision", that is, the data label is valid communication; if neither matches, then the state is marked as "not triggering a decision", that is, the data label is "invalid communication"; if only a partial match is found, then the state is marked as "potentially triggering a decision", that is, the data label is "potentially valid communication".

[0050] Optionally, the present invention can also perform coarse matching based on message type. The matching rule is as follows: when the decision type corresponding to the decision generated by the agent belongs to the set of decision types associated with the current input message type in the mapping table, it is determined that the decision and the message type have reached a clear match, and the triggering state is directly marked as "clearly triggered decision", that is, effective communication; if there is no match, the state is marked as "not triggered decision", that is, "invalid communication"; if there is only a partial match, the state is marked as "may trigger decision", that is, "potentially effective communication".

[0051] For example, if the input message type is an intent or request, the matching decision type is path planning or task assignment; if the input message type is a state update, the matching decision type is state adjustment.

[0052] For the same message, multiple matching rules can be applied, with the following priority: structured field matching > unstructured keyword matching > type matching. The result corresponding to the highest matching level is used as the final label of the communication data to ensure accuracy.

[0053] Based on this, each piece of data mentioned above has obtained a corresponding data tag. Within a preset time window, the number of effective communications, the number of potential effective communications, and the number of invalid communications are counted, and indicators such as effective communication rate, invalid communication rate, and potential effective communication rate are calculated respectively. Corresponding thresholds are set for each indicator. An alarm is triggered when the maximum threshold is exceeded or the minimum threshold is lowered. For example, when the invalid communication rate is as high as 50%, relevant intelligent agent data is sent to the anomaly feedback module of this invention for further processing.

[0054] Therefore, by employing a hierarchical matching strategy that combines structured field matching with unstructured keyword association, the actual triggering effect of communication on decision-making can be accurately identified within a time window, achieving a fine-grained, highly robust, and interpretable quantitative assessment of the effectiveness of communication.

[0055] According to one embodiment of the present invention, calculating the decision synchronization degree and decision stability of an agent based on a decision time sequence includes: extracting decision data of the agent at the timestamps in the decision time sequence, determining the effective decision agents in the decision data, and determining the target agent among the effective decision agents; calculating the decision synchronization degree of the agent at a single time node based on the target agent; traversing the decision time sequence of the agent in a preset time window and setting decision change rules; obtaining the number of changes between two adjacent decisions in the decision time sequence, and calculating the decision change frequency of the agent based on the number of changes and the preset time window, so as to calculate the decision stability of the agent based on the decision change frequency.

[0056] Specifically, the evaluation methods of related technologies often analyze the rationality of the decisions of individual agents while ignoring the synergy of decisions among multiple agents. This leads to the problem that individual agent decisions may be reasonable, but the overall decision-making fails. To identify this problem, a decision-making time sequence is constructed, and a comprehensive score of the collaborative stability among multiple agents is calculated, thereby accurately determining whether multiple agents have reached a "consensus and stable" collaborative state.

[0057] Specifically, such as Figure 3 As shown, a decision time series is constructed based on the collected agent decision data. The specific data structure includes {Agent ID: [Decision 1 (Decision content, timestamp t1, task stage), Decision 2 (Decision content, timestamp t2, task stage), ..., Decision n (Decision content, timestamp tn, task stage)], ...}. Based on the definition of the agent decision time series, the following key indicators are calculated.

[0058] First, the decision synchronization degree of agents is calculated based on the decision time sequence to measure whether multiple agents reach a consensus on the same goal at the same time point. For each time point t in the decision time sequence, the decision data of all agents at that time point is extracted. After removing agents with invalid decisions (such as those that did not generate a decision or whose decision content is empty), the effective decision agents in the decision data are determined, and the number of effective decision agents is counted, that is, the number of agents whose decisions are consistent with the common goal. The target agent is obtained, and the synchronization degree S of a single time point is calculated. To ensure the effectiveness and representativeness of the indicator, if the number of effective decision agents is less than 50% of the total number, it means that more than half of the agents cannot provide effective decisions, the data lacks group representativeness and has low reliability. In this case, it will interfere with the overall evaluation results. Therefore, the synchronization degree of that time point is marked as invalid and removed in subsequent calculations. The definition of the synchronization degree of a single time point is: S = Number of agents whose decisions align with the objective / Total number of agents making effective decisions; Simultaneously, the synchronization degree of all valid time nodes within the time window is arithmetically averaged to obtain the average synchronization degree within the window. The value range is [0,1]. The closer the value is to 1, the higher the degree of consensus in decision-making within that time window.

[0059] Secondly, computational decision stability is used to assess the consistency of multi-agent decision changes, avoiding collaborative disorder caused by frequent changes in the decisions of some agents.

[0060] Specifically, for each agent, its decision sequence within a preset time window is traversed, and a decision change judgment rule is set. When the core actions or key parameters of two adjacent decisions undergo substantial changes (such as changing from "go to area A" to "go to area B", or a change in movement speed parameter ≥ 20%), it is judged as a decision change. The number of decision changes made by the agent within the window is counted, and the agent's decision change frequency C is calculated. The expression for the decision change frequency C is: C = Number of decision changes / Duration of time window (unit: times / s); Next, calculate the standard deviation of this set of data based on the C-values ​​of all agents. This standard deviation reflects the dispersion of the frequency of decision changes among agents. The smaller the value, the more consistent the decision-making fluctuation rhythm of each agent is, and the stronger the group's coordination. If the change frequency of a certain agent exceeds three times the standard deviation of the average change frequency of all agents, it indicates that the decision-making fluctuation rhythm of that agent deviates greatly from the group, which is an extreme anomaly. It should be removed during the calculation to avoid interfering with the calculation results of the consistency of group fluctuation.

[0061] Among them, the maximum allowable fluctuation frequency is set. (Default is 5 times / s, which can be adjusted according to the task type, such as setting it to 2 times / s for high-precision collaborative tasks). Stability score converted to the [0, 1] interval The calculation formula is:

[0062] when hour, =0 indicates that the group's decision-making fluctuates wildly and has extremely poor stability; when hour, =1 indicates that all agent decisions remain unchanged, resulting in optimal stability.

[0063] in, This reflects the "degree of consensus". This reflects whether the consensus can be stably maintained; using either alone can lead to misjudgment and fail to meet actual collaborative needs. Therefore, with... and Subsequently, this invention uses a weighted summation method to calculate a comprehensive score for cooperative stability. This yields an indicator that balances consensus and stability, with the weighting of the indicator based on the principle of "consensus taking precedence over stability," as shown in the specific formula. =0.6× +0.4× Set a cooperative stability threshold. (Configurable, default 0.7, can be dynamically optimized based on historical data), when ≥ When, it is determined to be stable through collaborative decision-making; when < If the collaborative decision-making is unstable, the threshold can be dynamically adjusted according to the task stage to improve the accuracy of the assessment. For example, if the collaborative goal is not yet clear in the task initiation stage, it can be lowered to 0.5, and restored to 0.7 in the task execution stage.

[0064] In addition, unstable types can be further subdivided for subsequent anomaly feedback, but the judgment threshold needs to be subdivided according to the task type: For example, when <0.5 and When the value is greater than 2, it is judged as "instability of each agent swinging", that is, the multiple agents have no unified decision consensus and the decision fluctuation rhythm is completely inconsistent. Typical scenarios are multiple agents having conflicting goals and excessive communication delays leading to asynchronous decision information.

[0065] when ≥0.5 but When the value is greater than 2, it is judged as "short-term synchronization but violent fluctuation" type of instability. That is, multiple agents can reach a consensus on decision-making at the same time point, but the decision changes frequently, which makes it impossible to maintain the collaborative state. Typical scenarios are frequent changes in environmental perception data and insufficient adaptability of collaborative decision-making algorithms to environmental changes. For example, in scenarios with many dynamic obstacles, although multiple agents can adjust the path synchronously, the path decision changes frequently due to the frequent movement of obstacles.

[0066] Finally, if the agent's collaboration is deemed unstable based on the comprehensive score of collaborative stability, an anomaly feedback module needs to be reported. The reference type can be either "individual swinging" stable or "short-term synchronization but drastic fluctuations" unstable, depending on the actual situation.

[0067] Thus, by quantifying the instantaneous consensus strength (synchronicity) and behavioral evolution consistency (stability) of group decision-making in the temporal dimension, a dynamic, fine-grained, and interpretable stability assessment of the quality of multi-agent collaborative processes is achieved.

[0068] According to one embodiment of the present invention, obtaining the current round's profit deviation and current round's resource deviation of an agent includes: calculating the agent's average profit and average resource value based on a preset current time window; calculating the agent's profit deviation within the preset current time window based on the average profit value to obtain the current round's profit deviation; and calculating the agent's resource deviation within the preset current time window based on the average resource value to obtain the current round's resource deviation.

[0069] According to one embodiment of the present invention, the magnitude of profit change is obtained based on the profit deviation of the current round and the profit deviation of the previous round, and the magnitude of resource change is obtained based on the resource deviation of the current round and the resource deviation of the previous round, comprising: calculating a first difference between the profit deviation of a preset current time window and a preset previous time window to obtain the magnitude of profit change of the agent; and calculating a second difference between the resource deviation of a preset current time window and a preset previous time window to obtain the magnitude of resource change of the agent.

[0070] The preset current time window can be selected by those skilled in the art according to actual needs, and no specific limitation is made here.

[0071] Specifically, in the assessment of collaborative fairness, this invention breaks through the limitations of the traditional "static accounting after the task ends" and innovatively proposes a real-time fairness perception mechanism based on the dynamic deviation of rounds and its changing trend. This mechanism takes each round of tasks as the basic evaluation unit, quantifies the relative deviation of individual agents in terms of benefit acquisition and resource input, and further analyzes its cross-round evolution trend, thereby achieving early identification and accurate warning of typical problems such as "imbalance in allocation mechanism" or "external disturbances causing fairness deterioration".

[0072] Specifically, the fairness monitoring module identifies three types of fairness issues in the collaborative process by calculating the changes in agent rewards and resources within adjacent windows: persistently low individual rewards, unstable fairness, and deteriorating fairness. This module can identify anomalies after each round of task data collection and can issue real-time warnings without waiting for the entire task to end. It monitors the fairness and stability of the collaborative process during the process, rather than performing post-event fairness calculations as in traditional technologies.

[0073] In this embodiment of the invention, two types of core data of the agent need to be monitored in each round of tasks, including benefits and resources. Benefits are the "task result allocation weight" obtained by the agent after completing the task (such as the points after the completion of the sub-task and the task contribution score), and resources are the system-allocated resources obtained by the agent during the execution of the task (such as the computing resource quota and the communication bandwidth ratio).

[0074] Furthermore, the preset time window period is set as the execution cycle of each round of tasks, and therefore the window period setting will vary depending on the task type. For example, for time-driven tasks, the task is executed according to a fixed time period and there is no clear number of subtasks, so it needs to be divided into rounds according to time, such as 10 seconds; for tasks with fixed subtasks, the completion of all subtasks needs to be verified through task acceptance rules. If all agents report the "completed" status, then one round is considered to be completed.

[0075] The preset current time window for each round of tasks Calculate the average payoff of all agents in the swarm. and resource average , can be represented as:

[0076] Where n is the total number of agents within the time window. For agent i in the window The profit value, This refers to the amount of resources acquired.

[0077] To calculate the difference between each agent and the group mean, for each agent i, its value within a preset time window is calculated. Within the deviation of returns and resource deviation :

[0078] For each agent i, compare the current time window. Compared to the previous round time window The first difference in the deviation of the agent's profit is used to obtain the magnitude of the change in the agent's profit. ; Calculate the current time window Compared to the previous round time window The second difference in resource deviation is used to obtain the magnitude of resource change for the agent. The formula can be expressed as:

[0079] Specifically, when the reward value of an agent i satisfies the following conditions in k consecutive windows (k can be customized and is the number of task rounds, such as 10), the reward value is determined by the condition that the reward value of the agent i satisfies the following conditions. ( It is customizable and is a proportional factor. For example, setting it to 80% represents 80% of the average value, and the resource acquisition volume during the same period continuously meets the requirements. When resources are obtained at a level no lower than the average, but the returns are low, it indicates an unfair anomaly where individual returns are consistently low.

[0080] When the gains and resource changes of a certain intelligent agent i ( (This is a custom value representing the upper limit of the change range, such as 40%), and the percentage of agents exhibiting this phenomenon within a single window exceeds [a certain threshold]. hour( (This is a custom setting, representing the maximum percentage, such as 40%), which indicates that fairness is unstable.

[0081] When a communication interruption event occurs, compare it with the window before the interruption. and a window after the interruption After calculating the deviation of each agent, the average deviation of the entire system is calculated:

[0082] If the system average deviation is greater than (Custom value, representing the upper limit of average deviation, such as 25%, which means that the average deviation increases by more than 25% after the interruption), indicates that fairness has deteriorated during the agent collaboration process.

[0083] When the system indicates an unfair state according to the above definition, the specific exception is reported to the exception feedback module.

[0084] Therefore, by dynamically calculating the round-level deviation of individuals in terms of benefits and resources and the magnitude of changes across rounds, real-time, quantitative, and trend-aware monitoring of the fairness of multi-agent collaboration is achieved, effectively identifying typical fairness anomalies such as long-term unfairness, drastic fluctuations, or continuous deterioration.

[0085] In step S103, in response to any one of the evaluation results of communication effectiveness evaluation, collaborative decision stability evaluation, and collaborative fairness state evaluation being an abnormal result, cross-validation is performed using communication data, decision data, and state data to obtain cross-validation results. The cause of the agent's abnormality is obtained based on the cross-validation results, and a performance evaluation report of the agent is obtained based on the evaluation results and the cause of the abnormality.

[0086] According to one embodiment of the present invention, cross-validation is performed using communication data, decision data, and state data to obtain cross-validation results, and the cause of anomalies of the agent is obtained based on the cross-validation results. This includes: constructing a mapping table of communication effectiveness assessment, collaborative decision stability assessment, and collaborative fairness state assessment, as well as the corresponding cause of anomalies; and determining the cause of anomalies of the agent based on the priority of the communication effectiveness assessment, collaborative decision stability assessment, and collaborative fairness state assessment according to the mapping table.

[0087] Specifically, after independently evaluating three dimensions—communication effectiveness, collaborative decision-making stability, and collaborative fairness—this invention further constructs a root cause diagnosis engine based on multi-source data cross-validation. To this end, when the system detects any anomaly in the evaluation result, it immediately initiates the cross-validation process. By fusing communication data, decision data, and status data, and combining them with preset causal reasoning rules, it accurately locates the cause of the anomaly and provides a corresponding optimization strategy based on the cause and the evaluation result.

[0088] Specifically, such as Figure 4 As shown, the anomaly feedback module receives anomaly results from three modules: communication effectiveness assessment, collaborative decision-making stability assessment, and fairness status monitoring. It handles anomalies using a method of "anomaly data collection - multi-module cross-validation - anomaly location and handling." The process is as follows: When any module reports an anomaly, the evaluation data of all modules are automatically aggregated. Through multi-module cross-validation, false judgments are eliminated, and the anomaly-related data is locked. Finally, it is determined that the root cause of the anomaly is a communication problem, a decision-making problem, or a fairness problem, thereby improving the efficiency of handling anomalies in multi-agent collaboration.

[0089] Specifically, such as Figure 4 As shown, it mainly includes (1) abnormal data collection. When any module triggers an abnormal warning (such as the effective communication rate being less than 50%, the comprehensive score of collaborative stability being less than 0.7, or the fair state being abnormal), the module starts the data collection process to summarize relevant data, which mainly includes the following: abnormal details (including abnormal type, occurrence timestamp, involved agent ID, key indicator values) and the original data (messages / decision / logs) of the abnormal period, as well as cross-module related data (such as the decision synchronization degree of the communication abnormal period and the effective communication rate of the fair abnormal period), to ensure that the data is fully traceable; (2) Multi-module cross-validation: Cross-validation is used to eliminate misjudgments by a single module, mainly including communication anomaly verification, decision anomaly verification, and fairness anomaly verification. Then, the cause of the agent's anomaly is obtained based on the cross-validation results. Among them, communication anomaly verification: if the communication effectiveness evaluation module reports an anomaly, the collaborative decision stability module is checked simultaneously. If the score If the frequency drops significantly and the abnormal timing aligns with the communication anomaly, the cause of the agent's anomaly is determined to be "communication problems leading to abnormal collaboration." If only communication effectiveness is low while decision-making and fairness are normal, the cause of the agent's anomaly is determined to be "message redundancy / invalid transmission." For decision-related anomaly verification: if the collaborative decision-making stability module reports an anomaly, the communication and fairness modules are checked simultaneously. If communication effectiveness and fairness are normal, the cause of the agent's anomaly is determined to be "insufficient adaptability of the decision-making algorithm / deviation of the target consensus." If accompanied by communication anomalies, the cause of the agent's anomaly is prioritized as "communication problems leading to abnormal collaboration," with the root cause being the communication problem. For fairness anomaly verification: if the fairness state monitoring module reports an anomaly, the communication and decision-making modules are checked simultaneously. If communication and decision-making are both normal, the cause of the agent's anomaly is determined to be "individual task load imbalance (e.g., unreasonable resource allocation)," "individual fairness imbalance (unreasonable benefit distribution)," or "mismatch between resource input and benefit." If accompanied by drastic fluctuations in decision-making stability, the cause of the agent's anomaly is determined to be "deterioration of fairness due to decision instability," with the root cause being the decision-making problem.

[0090] (3) Anomaly Location and Handling: Based on the cross-validation results, corresponding optimization strategies and suggestions are given for the relevant anomaly causes. The anomaly causes mainly include communication problems, decision-making problems, and fairness problems. Communication problems include: communication problems leading to collaboration anomalies and message redundancy / invalid transmission. Therefore, when the anomaly cause is a communication problem, the corresponding optimization strategy is to optimize the communication protocol (when communication is asynchronous, adopt the "message confirmation receipt + timeout retransmission" mechanism to reduce latency) and clean up redundant messages (based on communication validity labels, automatically filter "invalid communication" messages and retain valid / potentially valid messages). Decision-making problems include: decision instability leading to deterioration of fairness, insufficient adaptability of decision-making algorithms, and deviation of target consensus. Therefore, when the anomaly cause is a decision-making problem, the corresponding optimization strategy is to optimize the communication protocol (when communication is asynchronous, adopt the "message confirmation receipt + timeout retransmission" mechanism to reduce latency) and clean up redundant messages (based on communication validity labels, automatically filter "invalid communication" messages and retain valid / potentially valid messages). When dealing with policy issues, the corresponding optimization strategies include stable decision-making (such as adding multi-source verification to key perception data like target location and task progress to ensure data reliability before triggering decision adjustments), automatic fine-tuning of targets according to task stages, and algorithm adaptation optimization (dynamically adjusting decision generation logic based on environmental perception data to improve adaptability to complex environments). For fairness issues, such as individual task load imbalance, individual fairness imbalance, and mismatch between resource input and returns, the corresponding optimization strategies when the cause of the anomaly is fairness are task diversion (e.g., diverting 30% of the sub-tasks of the abnormal agent to individuals with lower loads), adjusting resource allocation weights, and establishing a safety net mechanism (temporarily prioritizing high-return sub-tasks for agents with high resource input but low returns to quickly level off fairness bias and prevent a decline in collaborative motivation).

[0091] Therefore, based on the cross-validation results given by the three modules of communication, decision-making, and fairness, a structured evaluation of the overall collaborative performance of the system is conducted. This structured evaluation can pinpoint the cause of the agent's anomaly and provide corresponding optimization strategies and performance evaluation reports based on the cause of the anomaly. For example, if the cause of the anomaly is communication failure, the performance evaluation report will indicate that the system's collaborative performance is severely limited, mainly because the communication efficiency is below the threshold (e.g., 50%), resulting in 70% of the agents making asynchronous decisions, and it is recommended to enable the retransmission mechanism.

[0092] (4) Output of results: Output an anomaly assessment report, which clarifies the problem type, the agent involved, the related module data and the handling suggestions, so as to facilitate subsequent optimization.

[0093] Therefore, by constructing a priority mapping table between multi-dimensional evaluation results and anomaly causes, and based on cross-validation of three types of data—communication, decision-making, and fairness—we have achieved rapid, accurate, and interpretable automatic localization of the root causes of anomalies in multi-agent systems.

[0094] In summary, this invention presents a multi-dimensional performance evaluation method based on multi-agent collaboration. It obtains anomaly feedback by jointly monitoring communication effectiveness, collaborative decision-making stability, and collaborative fairness. Upon receiving anomaly feedback, it automatically reverse-engineers the root cause of the anomaly and provides relevant handling suggestions, mainly including the following aspects: (1) For the communication effectiveness assessment module, the communication efficiency-related indicators in the communication process are calculated through "communication-decision correlation", and an early warning is issued when the effective communication rate is low; (2) For the collaborative decision-making stability assessment module: comprehensively assess the decision-making stability of the agent collaboration process from two dimensions: synchronization degree (group goal consensus degree) and stability (consensus stability degree), and issue an early warning when it is unstable; (3) Fairness status monitoring: From "post-event accounting" to "in-event dynamic monitoring", the benefits and resources of the agent are monitored in real time after each round of tasks, and the deviation degree and the magnitude of deviation change are output to detect abnormal problems such as long-term low individual benefits, unstable fairness and deterioration of fairness, and issue early warnings. (4) Closed-loop anomaly location mechanism: integrates cross-validation of data from the three modules of communication, decision-making, and fairness, locks the root cause of anomalies according to priority, provides handling suggestions, and outputs an evaluation report, which greatly improves the efficiency of anomaly handling.

[0095] The performance evaluation method for intelligent agents proposed in this invention obtains communication effectiveness evaluation results, collaborative decision-making stability evaluation results, and collaborative fairness state evaluation results for intelligent agents by acquiring communication data, decision-making data, and state data. In response to any of the above three evaluation results being an abnormal result, cross-validation is performed using communication data, decision-making data, and state data to obtain the cause of the abnormality and a performance evaluation report. This solves the problems of single-dimensional performance evaluation of multi-agent systems, reliance on manual diagnosis, and ambiguous root causes in related technologies. By integrating dynamic evaluation and cross-validation of the three dimensions of communication effectiveness, collaborative decision-making stability, and collaborative fairness, precise root cause localization of performance anomalies in multi-agent systems is achieved, improving the comprehensiveness of the evaluation and the accuracy of the diagnosis.

[0096] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0097] Embodiments of the present invention also provide a performance evaluation device for an intelligent agent.

[0098] Figure 5 This is a block diagram illustrating the performance evaluation device for an intelligent agent according to an embodiment of the present invention.

[0099] like Figure 5 As shown, the performance evaluation device 10 of the intelligent agent includes: a first acquisition module 100, a second acquisition module 200 and a third acquisition module 300.

[0100] The first acquisition module 100 is used to acquire the communication data, decision data and status data of the intelligent agent; The second acquisition module 200 is used to obtain the communication effectiveness evaluation result, collaborative decision stability evaluation result, and collaborative fairness state evaluation result of the agent based on communication data, decision data, and state data. The third acquisition module 300 is used to respond to any one of the evaluation results of communication effectiveness evaluation, collaborative decision stability evaluation, and collaborative fairness state evaluation as an abnormal result. Then, it uses communication data, decision data, and state data to perform cross-validation to obtain cross-validation results, obtains the cause of the agent's abnormality based on the cross-validation results, and obtains the agent's performance evaluation report based on the evaluation results and the cause of the abnormality.

[0101] According to one embodiment of the present invention, the second acquisition module 200 includes: The first acquisition unit is used to construct a time window associated with communication data and analyze whether the communication data triggers the corresponding decision data within the time window, so as to obtain the communication effectiveness evaluation result based on the decision result; The second acquisition unit is used to construct a decision time sequence based on the decision data, calculate the decision synchronization degree and decision stability of the agent based on the decision time sequence, and calculate the comprehensive score of the agent's collaborative stability using a weighted summation method based on the decision synchronization degree and decision stability, so as to obtain the collaborative decision stability evaluation result based on the comprehensive score of collaborative stability. The third acquisition unit is used to acquire the current round's profit deviation, current round's resource deviation, previous round's profit deviation, and previous round's resource deviation based on state data. It also obtains the profit change magnitude based on the current round's profit deviation and the previous round's profit deviation, as well as the resource change magnitude based on the current round's resource deviation and the previous round's resource deviation, and obtains the collaborative fairness state assessment result based on the profit change magnitude and resource change magnitude.

[0102] According to one embodiment of the present invention, the first acquisition module 100 includes: The first determining unit is used to determine the attribute information to be collected, wherein the attribute information to be collected includes at least one of message identifier, sending agent, receiving agent, message type, message content characteristics and message sending time; The query unit is used to collect data on the attribute information to be collected, obtain communication data, and use a preset decision state query interface to query the data of the receiving agent to obtain decision data. The fourth acquisition unit is used to acquire the task execution logs and task monitoring logs generated during the operation of the intelligent agent, so as to obtain status data based on the task execution logs and task monitoring logs.

[0103] According to one embodiment of the present invention, the first acquisition unit includes: The association subunit is used to associate the message content in the agent's communication data with the corresponding decision data based on a preset matching strategy if the communication data triggers the corresponding decision data, so as to obtain the data tags of the communication data respectively. Obtain sub-units to obtain communication effectiveness assessment results based on data tags.

[0104] According to one embodiment of the present invention, the associated subunit includes: The judgment component is used to determine whether structured data exists in the message content and decision data; The first acquisition sub-component is used to extract the first target field of the message content and the second target field of the decision data if structured data exists in the message content and decision data, and to detect the matching degree between the first target field and the second target field. Based on the matching degree, data tags of multiple communication data are obtained respectively. Otherwise, the association between the message content and the decision data is constructed based on a preset matching method. The identification sub-component is used to identify preset keywords in the message content and decision information in the decision data that corresponds to the preset keywords; The second acquisition sub-component is used to obtain the data tag of the communication data based on the preset keyword and the decision information corresponding to the preset keyword.

[0105] According to one embodiment of the present invention, the second acquisition unit includes: The first computational subunit is used to extract the decision data of the agent at the timestamp in the decision time sequence based on the decision time sequence, determine the effective decision agents in the decision data, determine the target agent among the effective decision agents, and calculate the decision synchronization degree of the agent at a single time node based on the target agent. The traversal sub-unit is used to traverse the decision sequence of the agent within a preset time window and set the decision change rules. The second calculation subunit is used to obtain the number of changes between two adjacent decisions in the decision time sequence, and calculate the decision change frequency of the agent based on the number of changes and a preset time window, so as to calculate the decision stability of the agent based on the decision change frequency.

[0106] According to one embodiment of the present invention, the third acquisition unit includes: The third calculation subunit is used to calculate the average revenue and average resources of the agent based on a preset current time window; The fourth calculation subunit is used to calculate the agent's profit deviation in the current time window based on the average profit, to obtain the profit deviation of the current round, and to calculate the agent's resource deviation in the current time window based on the average resource, to obtain the resource deviation of the current round.

[0107] According to one embodiment of the present invention, the third acquisition unit includes: The fifth calculation subunit is used to calculate the first difference between the profit deviation between the preset current time window and the preset previous time window, and to obtain the profit change range of the agent. The sixth calculation subunit is used to calculate the second difference between the resource deviation between the preset current time window and the preset previous time window, so as to obtain the resource change range of the agent.

[0108] According to one embodiment of the present invention, the third acquisition module 300 includes: The building unit is used to construct a mapping table for communication effectiveness assessment, collaborative decision-making stability assessment, and collaborative fairness status assessment, as well as the corresponding causes of anomalies; The second determining unit is used to determine the cause of anomalies of the agent based on the priority of communication effectiveness assessment, collaborative decision-making stability assessment, and collaborative fairness state assessment according to the mapping table.

[0109] In summary, the descriptions of the features in the embodiments corresponding to the performance evaluation device for intelligent agents can be found in the relevant descriptions of the embodiments corresponding to the performance evaluation method for intelligent agents, and will not be repeated here.

[0110] Embodiments of the present invention also provide an electronic device, which may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0111] When the processor 602 executes the program, it implements the performance evaluation method for the intelligent agent provided in the above embodiments.

[0112] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.

[0113] The memory 601 is used to store computer programs that can run on the processor 602.

[0114] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0115] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0116] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0117] Processor 602 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0118] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the performance evaluation method for intelligent agents when running.

[0119] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0120] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the performance evaluation method for an intelligent agent.

[0121] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0122] The performance evaluation method for an intelligent agent provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A method for evaluating the performance of an intelligent agent, characterized in that, include: Acquire communication data, decision data, and state data of the intelligent agent; Based on the communication data, the decision data, and the state data, the communication effectiveness evaluation result, the collaborative decision stability evaluation result, and the collaborative fairness state evaluation result of the agent are obtained; If any one of the communication effectiveness assessment results, the collaborative decision stability assessment results, and the collaborative fairness state assessment results is an anomalous result, then cross-validation is performed using the communication data, the decision data, and the state data to obtain a cross-validation result. Based on the cross-validation result, the cause of the anomalousness of the agent is obtained, and based on the assessment result and the cause of the anomalousness, a performance assessment report of the agent is obtained.

2. The method according to claim 1, characterized in that, The process of obtaining the communication effectiveness evaluation result, collaborative decision stability evaluation result, and collaborative fairness state evaluation result of the agent based on the communication data, the decision data, and the state data includes: Construct a time window associated with the communication data, and analyze whether the communication data triggers corresponding decision data within the time window, so as to obtain the communication effectiveness evaluation result based on the decision result; A decision time series is constructed based on the decision data. The decision synchronization degree and decision stability of the agent are calculated based on the decision time series. Based on the decision synchronization degree and the decision stability, a comprehensive score of the agent's collaborative stability is calculated using a weighted summation method. The collaborative decision stability evaluation result is obtained based on the comprehensive score of collaborative stability. Based on the state data, the agent's current round profit deviation, current round resource deviation, previous round profit deviation, and previous round resource deviation are obtained. The profit change magnitude is obtained based on the current round profit deviation and the previous round profit deviation, and the resource change magnitude is obtained based on the current round resource deviation and the previous round resource deviation. Finally, the cooperation fairness state assessment result is obtained based on the profit change magnitude and the resource change magnitude.

3. The method according to claim 1, characterized in that, The acquisition of the agent's communication data, decision data, and state data includes: The attribute information to be collected is determined, wherein the attribute information to be collected includes at least one of message identifier, sending agent, receiving agent, message type, message content characteristics and message sending time; Data is collected from the attribute information to be collected to obtain the communication data, and the decision data is obtained by querying the data of the receiving agent using a preset decision state query interface. Obtain task execution logs and task monitoring logs generated during the operation of the intelligent agent, and obtain the status data based on the task execution logs and the task monitoring logs.

4. The method according to claim 2 or 3, characterized in that, Within a time window, analyze whether the communication data triggers corresponding decision data, and obtain the communication effectiveness evaluation result based on the decision result, including: If the communication data triggers the corresponding decision data, then based on the preset matching strategy, the message content in the agent's communication data is associated with the corresponding decision data to obtain the data tags of the communication data respectively. The communication effectiveness assessment result is obtained based on the data tags.

5. The method according to claim 4, characterized in that, The method, based on a preset matching strategy, associates the message content in the agent's communication data with the corresponding decision data to obtain data tags for the communication data, including: Determine whether structured data exists in the message content and the decision data; If the structured data exists in the message content and the decision data, then the first target field of the message content and the second target field of the decision data are extracted, and the matching degree between the first target field and the second target field is detected. Based on the matching degree, data tags for multiple communication data are obtained respectively. Otherwise, the association between the message content and the decision data is constructed based on a preset matching method. Identify preset keywords in the message content and decision information corresponding to the preset keywords in the decision data; The data tag of the communication data is obtained based on the preset keyword and the decision information corresponding to the preset keyword.

6. The method according to claim 2, characterized in that, The calculation of the agent's decision synchronization degree and decision stability based on the decision time sequence includes: Based on the decision time sequence, the decision data of the agent at the timestamp in the decision time sequence is extracted, and the effective decision agents in the decision data are determined. The target agent is determined among the effective decision agents. Based on the target agent, the decision synchronization degree of the agent at a single time node is calculated. The decision sequence of the agent within a preset time window is traversed one by one, and decision change rules are set. The number of changes between two adjacent decisions in the decision time sequence is obtained, and the decision change frequency of the agent is calculated based on the number of changes and the preset time window, so as to calculate the decision stability of the agent based on the decision change frequency.

7. The method according to claim 2, characterized in that, The process of obtaining the agent's current round profit deviation and current round resource deviation includes: Based on a preset current time window, calculate the average revenue and average resources of the agent; The agent's profit deviation in the current time window is calculated based on the average profit, to obtain the profit deviation of the current round; and the agent's resource deviation in the current time window is calculated based on the average resource, to obtain the resource deviation of the current round.

8. The method according to claim 2 or 7, characterized in that, The step of obtaining the profit change range based on the profit deviation of the current round and the profit deviation of the previous round, and obtaining the resource change range based on the resource deviation of the current round and the resource deviation of the previous round, includes: Calculate the first difference between the profit deviation of the preset current time window and the preset previous time window to obtain the profit change range of the agent; Calculate the second difference between the resource deviation of the preset current time window and the preset previous time window to obtain the resource change range of the agent.

9. The method according to claim 1, characterized in that, The process of using the communication data, the decision data, and the state data to perform cross-validation to obtain cross-validation results, and then determining the cause of the agent's anomaly based on the cross-validation results, includes: Construct a mapping table for communication effectiveness assessment, collaborative decision-making stability assessment, and collaborative fairness status assessment, as well as the corresponding causes of anomalies; Based on the mapping table, the cause of the agent's anomaly is determined according to the priority of the communication effectiveness assessment, the collaborative decision-making stability assessment, and the collaborative fairness state assessment.

10. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the performance evaluation method for an intelligent agent as described in any one of claims 1 to 9.