Electric power large model agent operation evaluation method, device, equipment and medium
By generating task flows, analyzing the behavioral characteristics and effectiveness of intelligent agents, the dynamic adaptability of intelligent agents in the power large model is evaluated, solving the problem that traditional evaluation methods cannot capture their coordination in dynamic environments, and realizing a comprehensive evaluation of them in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to accurately capture the adaptive adjustment capabilities of large-scale power model agents in dynamic environments, and traditional evaluation methods cannot effectively characterize their dynamic coordination in complex environments.
By generating task flows, extracting agent behavior records, analyzing coordination behavior characteristics, quantifying coordination effectiveness, generating a comprehensive coordination evaluation report, assisting in evaluating the agent's dynamic adaptability, identifying policy adjustment events and environmental change points, calculating behavior adjustment delays, evaluating policy adjustment effects, and generating a comprehensive coordination evaluation report.
It enables real-time evaluation of an agent's coordination capabilities and strategy adjustments in dynamic environments, penetrating behavioral appearances to quantify its internal coordination mechanisms and assess the effectiveness of its decision-making logic and strategy adjustments in the face of sudden tasks and resource constraints.
Smart Images

Figure CN121787728A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence diagnosis, and in particular relates to a method, device, equipment and medium for evaluating the operation of a large-scale power intelligent agent. Background Technology
[0002] With the deepening of the intelligent transformation of power systems, the technology of large-scale power model intelligent agents has emerged. Its core characteristics are the ability to process massive amounts of heterogeneous data and possess complex autonomous reasoning and decision-making capabilities to cope with uncertainties in scenarios such as power grid dispatching and fault handling. This makes evaluating its operational performance, especially its coordination capabilities in dynamic environments with multiple concurrent tasks and intense resource competition, crucial for ensuring its reliable application. Currently, the evaluation of intelligent agent capabilities mainly relies on the traditional approach combining static testing and post-event result analysis. Traditional techniques typically use pre-programmed test cases or replays of historical scenarios to verify the agent's functionality. The evaluation focus is on isolated metrics such as the final accuracy of task processing and response time, or on testing its logic under simplified, fixed task sequences. This method measures the agent's behavior as static output in a deterministic environment. However, current evaluation methods are insufficient to effectively characterize the dynamic coordination exhibited by intelligent agents in real, complex environments. Faced with rapidly changing power grid conditions and a multitude of intertwined tasks, how do intelligent agents understand, weigh, and allocate their internal attention and computing resources, and how do their decision-making logic adapt to environmental evolution? These dynamic, continuous, and internally coupled behavioral characteristics are difficult to accurately capture and deeply diagnose within the traditional static and fragmented evaluation framework. Summary of the Invention
[0003] Based on this, it is necessary to provide a method, apparatus, equipment, and medium for evaluating the operation of a large-scale power model intelligent agent that can accurately capture how the decision-making logic of the diagnostic agent adapts to environmental evolution, in order to address the aforementioned technical problems.
[0004] Firstly, this application provides a method for evaluating the operation of a large-scale power model intelligent agent, including:
[0005] Based on the evaluation purpose, a task flow is generated and input into the agent to be evaluated to obtain the agent's behavior record;
[0006] Feature extraction is performed on the agent's behavior records to obtain a coordination behavior feature dataset. Based on the coordination behavior feature dataset, the coordination strategy of the agent to be evaluated is inferred to obtain the agent's coordination strategy.
[0007] Based on the dataset of coordination behavior characteristics, indicators of coordination effectiveness are quantified to obtain a numerical table of coordination effectiveness.
[0008] Based on the agent coordination strategy and coordination effectiveness numerical table, the dynamic adaptability of the agent to be evaluated is analyzed, and a dynamic adaptability analysis report is obtained.
[0009] Based on the dynamic adaptation analysis report and the coordination effectiveness numerical table, a comprehensive coordination evaluation report is generated; the comprehensive coordination evaluation report is used to assist in evaluating the current capabilities of the agent to be evaluated.
[0010] Furthermore, based on the agent's coordination strategy and coordination effectiveness numerical table, the dynamic adaptability of the agent to be evaluated is analyzed, resulting in a dynamic adaptability analysis report, including:
[0011] Environmental change events are extracted from the task flow, and based on the agent coordination strategy, the time points when the agent to be evaluated adjusts its strategy are identified to obtain the behavior change points; the environmental change events include at least one of the following: changes in task attributes, changes in system state, and changes in resource constraints.
[0012] By matching environmental change events and behavioral change points in time, a response matching table is obtained;
[0013] Based on the response pairing table, the time difference between each environmental change event and behavioral change point is calculated to obtain the behavioral adjustment delay measurement value.
[0014] Based on the coordination effectiveness numerical table and behavioral change points, the dynamic response effect of strategy adjustment events is judged, and the strategy adjustment effect evaluation is obtained.
[0015] Based on the evaluation of the effect of policy adjustment and the measurement of the delay in behavior adjustment, the overall adaptability of the agent under evaluation is assessed, and a dynamic adaptation analysis report is obtained.
[0016] Furthermore, based on the coordination effectiveness numerical table and behavioral change points, the dynamic response effect of strategy adjustment events is judged to obtain an evaluation of the strategy adjustment effect, including:
[0017] Define analysis windows of fixed length before and after the behavior change point, and extract all coordination effectiveness indicators within the analysis window based on the coordination effectiveness numerical table to obtain effectiveness indicator data.
[0018] Calculate the statistical measures of the performance indicators to obtain the performance indicator statistical measures table; based on the points of behavioral change, compare and analyze the statistical measures of the indicators between the windows to obtain the before-and-after difference comparison table.
[0019] Based on statistical testing methods, statistical significance tests were performed on the performance indicator data to obtain the significance test results;
[0020] Based on the before-and-after difference comparison table and the significance test results, the adjustment effect of the behavioral change points is mapped to the corresponding level to obtain the effect direction judgment table.
[0021] Based on the effect direction judgment table, an evaluation of the effect of strategy adjustment is generated.
[0022] Furthermore, based on the evaluation of the policy adjustment effect and the measurement of the behavior adjustment delay, the overall adaptability of the agent under evaluation is assessed, resulting in a dynamic adaptation analysis report, including:
[0023] Based on the points of behavioral change, the data of the strategy adjustment effect evaluation and the behavioral adjustment delay measurement are combined to obtain a comprehensive strategy adjustment evaluation table.
[0024] Based on the comprehensive evaluation table of strategy adjustments, the average adjustment latency and adjustment success rate are calculated using the following formulas to obtain a summary of adaptive performance indicators:
[0025]
[0026]
[0027] in, To represent the average adjustment delay, n is the total number of behavior change points, and i is the index of the behavior change point. The adjustment delay time is the time for the i-th behavior change point. To adjust the success rate, Classify the adjustment effect at the i-th behavior change point. For indicator functions;
[0028] Based on the comprehensive evaluation table of strategy adjustment, strategy adjustment events are grouped to obtain clustering results of adaptive behavior patterns.
[0029] Based on the summary of adaptive performance indicators and the clustering results of adaptive behavior patterns, the overall adaptive capability of the agent to be evaluated is summarized, and a dynamic adaptive analysis report is obtained.
[0030] Furthermore, before generating a comprehensive coordination assessment report based on the dynamic adaptation analysis report and coordination effectiveness numerical table, the following steps are also included:
[0031] Traverse the coordination efficiency value table, filter out the timestamps corresponding to the indicators of coordination efficiency that are lower than the predefined industry benchmark value, and obtain the list of inefficient times;
[0032] Based on the list of inefficient times, a subset of environmental features corresponding to each timestamp is extracted from the coordinated behavior feature dataset to obtain inefficient environmental features.
[0033] Inefficient environment characteristics are input into a predefined expert rule base to obtain standard response strategies. The differences between the agent coordination strategies corresponding to inefficient environment characteristics and the standard response strategies are compared to obtain a behavior difference analysis matrix.
[0034] Based on behavioral difference analysis, the reasoning process of the agent under evaluation during decision-making is analyzed to identify decision logic defects.
[0035] Based on decision-making logic analysis, a coordination capacity bottleneck report is obtained; the coordination capacity bottleneck report is used to assist in generating a comprehensive coordination assessment report.
[0036] Furthermore, by comparing the differences between the agent's coordination strategy and the standard response strategy corresponding to the characteristics of inefficient environments, a behavioral difference analysis matrix is obtained, including:
[0037] The agent coordination strategy and the standard response strategy are matched and aligned to obtain an alignment pairing table;
[0038] Based on the alignment and pairing table, the differences between the agent coordination strategy and the standard response strategy are compared to obtain a list of decision difference points;
[0039] Based on the list of decision-making discrepancies, calculate the overall discrepancy index to obtain a summary of the overall discrepancy index; the overall discrepancy index includes at least one of the following: decision consistency rate, response time deviation, and critical decision error rate.
[0040] Based on the summary of overall difference indicators, a behavioral difference analysis matrix is obtained.
[0041] Secondly, this application also provides a power large-scale model intelligent agent operation evaluation device, comprising:
[0042] The recording module is used to generate a task flow based on the evaluation purpose, and input the task flow into the agent to be evaluated to obtain the agent's behavior record;
[0043] The feature module is used to extract features from the agent's behavior records to obtain a coordination behavior feature dataset, and based on the coordination behavior feature dataset, to infer the coordination strategy of the agent to be evaluated, thus obtaining the agent's coordination strategy.
[0044] The quantification module is used to quantify the indicators of coordination effectiveness based on the coordination behavior feature dataset, and obtain a numerical table of coordination effectiveness.
[0045] The dynamic module is used to analyze the dynamic adaptability of the agent under evaluation based on the agent's coordination strategy and coordination performance numerical table, and to obtain a dynamic adaptability analysis report.
[0046] The reporting module is used to generate a comprehensive coordination assessment report based on the dynamic adaptation analysis report and the coordination effectiveness numerical table; the comprehensive coordination assessment report is used to assist in evaluating the current capabilities of the agent to be evaluated.
[0047] Thirdly, this application also provides a computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement any step of the method provided in the first aspect of this application.
[0048] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of the method provided in the first aspect of this application.
[0049] The aforementioned method, apparatus, equipment, and medium for evaluating the operation of a large-scale power model intelligent agent, based on the evaluation objective, generates a task flow and inputs it into the intelligent agent to be evaluated, obtaining the agent's behavior records. Features are extracted from the agent's behavior records to obtain a coordination behavior feature dataset. Based on this dataset, the coordination strategy of the agent to be evaluated is inferred, resulting in the agent's coordination strategy. Based on the coordination behavior feature dataset, indicators of coordination effectiveness are quantified, resulting in a coordination effectiveness numerical table. Based on the agent's coordination strategy and the coordination effectiveness numerical table, the dynamic adaptability of the agent to be evaluated is analyzed, resulting in a dynamic adaptation analysis report. Based on the dynamic adaptation analysis report and the coordination effectiveness numerical table, a comprehensive coordination evaluation report is generated. This comprehensive evaluation report is used to assist in evaluating the current capabilities of the agent to be evaluated. It can penetrate the agent's behavioral appearance, quantify the dynamic adaptability of its internal coordination mechanism, capture agent behavioral characteristics in real time, quantify coordination performance, and effectively evaluate whether the agent's decision-making logic is reasonable during sudden tasks and whether its strategy adjustments are effective when resources are scarce. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 This is a schematic diagram of the process of a power large-scale model intelligent agent operation evaluation method provided in an embodiment of the present invention;
[0052] Figure 2 This is a schematic diagram of the structure of a power large-scale intelligent agent operation evaluation device provided in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0054] In one embodiment, such as Figure 1 As shown, a method for evaluating the operation of a large-scale power system intelligent agent is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0055] Step 101: Based on the evaluation purpose, generate a task flow and input the task flow into the agent to be evaluated to obtain the agent's behavior record.
[0056] The evaluation objective refers to the specific goals or dimensions of the agent's coordination capabilities that this evaluation aims to examine and verify. For example, it may test the agent's cooperative ability under resource competition, its reorganization ability under sudden failures, or its negotiation ability under incomplete information, thus determining the design direction of the task flow. The task flow refers to a pre-designed, ordered, or branching sequence of task scenarios, environmental states, and events, simulating various situations where the agent may need to coordinate with other agents or the environment. It is a set of test cases used to stimulate and observe the agent's coordination behavior. The agent to be evaluated refers to the agent with decision-making and behavioral capabilities being evaluated; it is the object of the evaluation. The agent behavior record refers to the complete raw data log recorded after the task flow is input into the agent to be evaluated. It includes all observable outputs generated by the agent in responding to the task flow, such as messages sent, actions performed, decisions made, internal states, and corresponding timestamps and environmental state snapshots. Based on a clear evaluation objective, the terminal designs and generates a task flow that can cover a variety of coordination challenges and scenarios. It runs this task flow in a simulation environment and submits the environmental state, task requirements, and behavior of other agents defined in each step of the task flow as input to the agent to be evaluated. Simultaneously, it records all the agent's reactions and outputs to the environmental inputs, forming a detailed log arranged in chronological order.
[0057] Step 102: Extract features from the agent's behavior records to obtain a coordination behavior feature dataset, and infer the coordination strategy of the agent to be evaluated based on the coordination behavior feature dataset to obtain the agent's coordination strategy.
[0058] The coordination behavior feature dataset refers to a series of feature vectors or data sequences extracted from the raw and complex records of agent behavior through calculation and transformation. These features quantify the agent's coordination behavior and state, and are related to coordination ability. Features are quantitative standards used to characterize a specific dimension of coordination behavior; they are measures calculated from the raw behavior. For example, communication features may include: the number of collaborative messages sent / received per unit time, the distribution ratio of message types, and the number of communication objects; task behavior features may include: the proportion of proactively undertaking sub-tasks, the delay in task initiation and response, and the contribution to shared goals; resource interaction features may include: the frequency of resource requests / releases, the duration of resource sharing / occupancy, and the number of resource conflicts; and timing and synchronization features may include: the synchronization rate with other agent actions, the periodicity or rhythm of its own behavior, and the time variance of responses to environmental events. Agent coordination strategy refers to an abstract description of the inherent rules, decision-making patterns, or behavioral logic adopted by the agent during the coordination process, reflecting how the agent coordinates. The terminal performs feature extraction, traverses the agent's behavior records, defines a set of key feature indicators based on coordination theory, and generates time series or statistical values of these indicators from the original logs through calculation. For example, the number of cooperation request messages sent by the agent per minute is calculated as the active coordination frequency. All extracted features are aligned by time to form a structured coordination behavior feature dataset. Policy inference is then performed. Based on the feature dataset, pattern recognition, time series analysis, or machine learning methods are used to analyze the combination and change patterns of agent behavior features in different contexts, thereby inferring the potential and stable coordination decision-making patterns or strategies followed by the agent.
[0059] Step 103: Based on the coordination behavior feature dataset, quantify the indicators of coordination effectiveness to obtain a coordination effectiveness numerical table.
[0060] Specifically, coordination effectiveness indicators refer to specific evaluation standards used to measure the final effect and efficiency of coordination behavior. These indicators directly answer the question of how well the coordination is done. Examples include: task completion indicators such as overall task success rate, sub-task completion rate, and goal achievement time; efficiency and cost indicators such as the total communication resources consumed to complete a unit task, computing resource overhead, physical resource consumption, or total number of action steps; quality and robustness indicators such as the stability of task completion results, performance retention in the presence of interference or noise, and the average resource utilization rate of the system as a whole during the coordination process; and collaboration-specific indicators such as the success rate of conflict resolution, the balance between overall team benefits and individual benefits, and the redundancy or complementarity of collaborative actions. The coordination effectiveness numerical table refers to the specific numerical results corresponding to each coordination effectiveness indicator, existing in tabular or time series form, clearly indicating the agent's performance score on each effectiveness indicator in the evaluation. Based on the coordination behavior feature dataset and supplementary information obtained from the task flow and environment, the terminal calculates the value of each coordination effectiveness indicator according to a predefined effectiveness calculation formula. Alternatively, using the team's total revenue as a performance indicator requires summing the revenues of all agents from the feature dataset and task results, which is purely computational and quantitative.
[0061] Step 104: Based on the agent coordination strategy and coordination effectiveness numerical table, analyze the dynamic adaptability of the agent to be evaluated and obtain a dynamic adaptability analysis report.
[0062] Specifically, dynamic adaptability refers to the ability of an agent to actively or passively adjust its coordination strategy in response to changes in the environment, tasks, or partners within a task flow, in order to maintain or improve coordination efficiency. A dynamic adaptation analysis report is an analysis document that details when, why, and how the agent adjusted its strategy within the task flow, and whether this adjustment was timely and effective in improving efficiency. It includes response latency, evaluation of adjustment effects, and a summary of adaptation patterns. The terminal performs correlation analysis by combining the agent's coordination strategy and the coordination efficiency numerical table. The core of this analysis is correlation analysis, which identifies designed environmental change points or sudden interference events in the task flow, locates the time points where significant strategy changes occur on the agent's coordination strategy timeline, analyzes the temporal relationship between strategy change points and environmental change points, and compares the situation before and after the strategy change. The coordination efficiency numerical table is used to see if relevant efficiency indicators have significantly improved, deteriorated, or remained stable, comprehensively judging the agent's response speed, adjustment accuracy, and final effect to dynamic changes.
[0063] Step 105: Based on the dynamic adaptation analysis report and the coordination effectiveness numerical table, generate a comprehensive coordination evaluation report; the comprehensive coordination evaluation report is used to assist in evaluating the current capabilities of the agent to be evaluated.
[0064] The comprehensive coordination assessment report is the final, integrated assessment conclusion document. It integrates the analysis results from all steps, providing a comprehensive, structured, and conclusive evaluation of the agent's coordination capabilities. This may include quantifiable performance scores, strategy characteristic analysis, dynamic adaptability ratings, a summary of strengths and weaknesses, and improvement suggestions. The terminal summarizes, cross-validates, and weights the quantitative scores from the coordination performance table with the qualitative analysis conclusions from the dynamic adaptability analysis report. For example, combining high performance scores with strong adaptability evidence leads to a conclusion of strong coordination capabilities; if high performance is found but poor adaptability, it indicates excellent performance in stable environments but insufficient ability to cope with sudden changes. The report summarizes the agent's overall coordination capabilities, main behavioral patterns, adaptability strength, and specific bottlenecks.
[0065] This embodiment provides a method for evaluating the operation of a large-scale power system agent. Based on the evaluation objective, a task flow is generated and input into the agent to be evaluated, resulting in agent behavior records. Features are extracted from these records to obtain a coordination behavior feature dataset. Based on this dataset, the coordination strategy of the agent is inferred, yielding the agent's coordination strategy. Based on the coordination behavior feature dataset, indicators of coordination effectiveness are quantified, resulting in a coordination effectiveness numerical table. Based on the agent's coordination strategy and the coordination effectiveness numerical table, the dynamic adaptability of the agent is analyzed, resulting in a dynamic adaptability analysis report. Based on the dynamic adaptability analysis report and the coordination effectiveness numerical table, a comprehensive coordination evaluation report is generated. This report assists in evaluating the current capabilities of the agent. Through these methods, the method can penetrate the agent's behavioral appearance, quantify the dynamic adaptability of the agent's internal coordination mechanism, capture agent behavioral characteristics in real time, quantify coordination performance, and effectively evaluate the rationality of the agent's decision-making logic in the event of emergencies or multi-task parallelism, and the effectiveness of strategy adjustments under resource constraints.
[0066] In one embodiment, based on the agent coordination strategy and coordination performance numerical table, the dynamic adaptability of the agent to be evaluated is analyzed to obtain a dynamic adaptability analysis report, including:
[0067] Step 201: Extract environmental change events from the task flow, and based on the agent coordination strategy, identify the time points when the agent to be evaluated adjusts its strategy to obtain the behavior change points; the environmental change events include at least one of the following: task attribute changes, system state changes, and resource constraint changes.
[0068] Among them, environmental change events refer to changes in external conditions that are actively introduced or occur naturally within a pre-designed task flow, posing a challenge to the agent's coordinated behavior or requiring a response. These include changes in task attributes, such as changes in task objectives, priority adjustments, or the insertion of new tasks; changes in system state, such as partner failures, sudden increases in communication network latency, loss of perceived information, and changes in resource constraints, such as reduced available computing resources, occupied shared physical resources, or limited energy budgets. It is a data record containing the event type, occurrence timestamp, and a description of the event's specific content. Policy adjustment events refer to the internal process by which the agent's decision-making logic or behavioral patterns significantly change when facing environmental changes; these are decision-level shifts inferred through behavioral pattern analysis. Behavioral change points refer to the specific time points at which the agent's policy adjustment events occur, determined after analysis. It is a list of timestamps, with each time point marking a recognizable shift in the agent's coordinated strategy. The terminal extracts environmental change events, replays the task flow design scripts or logs, locates and records all pre-set or system-triggered key environmental disturbances, forming a list of environmental change events with precise timestamps. It identifies the time points of policy adjustments, and based on the previously inferred agent coordination strategy, uses change point detection algorithms or rules to analyze the policy sequence, identifying the locations where significant, non-random transitions or turning points occur. These locations are marked as behavioral change points. This clarifies the externally applied stimulus line and the agent's internal response line, laying the foundation for subsequent analysis of how the agent responds to external changes.
[0069] Step 202: Match environmental change events and behavioral change points over time to obtain a response matching table.
[0070] Specifically, the response pairing table is a data structure that associates environmental change events with behavioral change points. It lists each event identified as a potential response to a specific environmental change and its associated behavioral change point, along with the time difference between them. This reflects a one-to-one or one-to-many correspondence and may also include unresponsive environmental events. The terminal iterates through each environmental change event, searching for behavioral change points within the time frame following its occurrence. If a behavioral change point occurs after an environmental change event, and the time interval between the two is reasonable, and there is a logical causal relationship between the event type and the content of the strategy change, then the event is paired with that behavioral change point. The matching process is based on a fixed time window, but can also be based on more complex causal inference logic. This establishes a hypothetical causal relationship between external stimuli and internal responses, enabling the assessment to focus on the interaction between environmental changes and strategy adjustments.
[0071] Step 203: Based on the response pairing table, calculate the time difference between each environmental change event and behavioral change point to obtain the behavioral adjustment delay measurement value.
[0072] Specifically, the behavior adjustment delay measurement refers to the time difference calculated for each pair of environmental change events and behavior change points in the response pairing table. It quantifies the time elapsed from the environmental change to the agent being detected and beginning to adjust its coordination strategy. It is a list or set containing the specific delay time for each response. For each pairing record in the response pairing table, the terminal reads the timestamp of the environmental change event and the timestamp of the behavior change point, subtracts the former from the latter to obtain the behavior adjustment delay time for a single event, traverses all pairing rows, calculates a series of delay values, and performs statistical analysis on these delay values, such as calculating the average, median, maximum, minimum, and standard deviation, to summarize the overall timeliness of the agent's response and quantify the agent's response speed.
[0073] Step 204: Based on the coordination effectiveness numerical table and behavioral change points, determine the dynamic response effect of the strategy adjustment event and obtain the strategy adjustment effect evaluation.
[0074] Specifically, strategy adjustment effectiveness evaluation refers to the qualitative or quantitative assessment of the effectiveness of each strategy adjustment event. It evaluates whether the agent's coordination efficiency improves, deteriorates, or remains unchanged after making a strategy adjustment. The evaluation results are labeled with the effect classification for each behavioral change point, including significant improvement, slight improvement, no impact, negative effect, and corresponding confidence level or strength of evidence. The terminal performs a longitudinal comparative analysis process, centering on each behavioral change point, defining a time period before and after that point as a pre-adjustment window and a post-adjustment window. Data on relevant efficiency indicators within these two windows are extracted from the coordination efficiency value table. Statistical methods are used to analyze the changes in indicators before and after the adjustment. If the efficiency indicators in the post-adjustment window show a statistically significant improvement compared to the pre-adjustment window, the strategy adjustment is considered effective; otherwise, it is considered ineffective or negative. Finally, an effect evaluation is assigned to each behavioral change point.
[0075] Step 205: Based on the evaluation of the policy adjustment effect and the measurement of the behavior adjustment delay, assess the overall adaptability of the agent to be evaluated and obtain a dynamic adaptation analysis report.
[0076] The overall adaptability refers to a comprehensive evaluation of the agent's ability to maintain or improve coordination performance in a dynamically changing environment, combining response speed and adjustment effectiveness. The dynamic adaptation analysis report is a comprehensive analytical conclusion document that integrates speed metrics and quality assessments, summarizing the agent's dynamic adaptability from multiple dimensions. The report includes, but is not limited to, average response latency, adjustment success rate, adaptation patterns to different types of environmental changes, response stability, and qualitative conclusions regarding the strength of overall adaptability. The terminal summarizes and cross-analyzes two sets of data: behavior adjustment latency measurement and strategy adjustment effectiveness evaluation. For example, it can calculate the average latency and the proportion of effective strategies for all adjustment events; analyze whether there are differences in the agent's latency and effectiveness under different types of environmental changes; and observe response patterns, such as immediate small-step rapid adjustments or large adjustments after a delay. Based on this comprehensive analysis, a holistic conclusion about the agent's adaptability is formed.
[0077] This embodiment provides a comprehensive and structured assessment of the agent's adaptability by integrating two dimensions: latency and effectiveness. The resulting dynamic adaptation analysis report clearly reveals the agent's strengths and weaknesses in dynamic environments.
[0078] In one embodiment, based on a coordination effectiveness table and behavioral change points, the dynamic response effect of strategy adjustment events is determined to obtain a strategy adjustment effect evaluation, including:
[0079] Step 301: Define analysis windows of fixed length before and after the behavior change points, and extract all coordination effectiveness indicators within the analysis windows based on the coordination effectiveness value table to obtain effectiveness indicator data.
[0080] The analysis window refers to two equal-length, continuous time periods defined before and after the point of behavioral change, using the point of behavioral change as the time reference origin. For example, with the time of the behavioral change as point 0, [-T, 0) is defined as the pre-adjustment window, and [0, +T) as the post-adjustment window, where T is a preset fixed duration. The two windows represent a period of time before and after the strategy adjustment, respectively. The performance indicator data refers to the raw numerical sequence of all performance indicators extracted from the coordination performance value table, specifically for the pre-adjustment and post-adjustment windows corresponding to each behavioral change point. It is a data slice arranged by time points, containing the specific values of each performance indicator at each moment within the window. For each identified behavioral change point, the terminal defines a time window. Using the timestamp of the behavioral change point as the boundary, it traces back a fixed duration T to determine the pre-adjustment window; and continues forward for the same fixed duration T to determine the post-adjustment window. It extracts data, accesses the data source storing the coordination performance value table, and accurately finds and extracts all performance indicator records falling within the window time range based on the start and end times of these two windows. For each behavioral change point, two sets of performance indicator data will be generated, representing the performance before and after the adjustment, respectively.
[0081] Step 302: Calculate the statistical measures of the performance indicators to obtain the performance indicator statistical measures table; based on the behavioral change points, compare and analyze the statistical measures of the indicators between the windows to obtain the before-and-after difference comparison table.
[0082] Specifically, indicator statistics refer to key numerical values used to summarize and describe the distribution characteristics of a set of data. Commonly used statistics include mean, median, standard deviation, maximum, and minimum. A performance indicator statistics table is a summary table compiled after calculating various statistics for performance indicator data within each analysis window, summarizing the overall level of each performance indicator within that window. A before-and-after difference comparison table is a structured list used for systematic comparison. For each point of behavioral change and each performance indicator, it lists the corresponding statistics for that indicator in the window before adjustment and the corresponding statistics in the window after adjustment, and calculates the absolute difference or relative rate of change between the two. The terminal performs preliminary data summarization and comparison, calculates statistics, and calculates the main statistics of each performance indicator in the window for each performance indicator data, forming a performance indicator statistics table. The tables are then compared side by side using the behavior change point as an index. The performance indicator statistics tables of the two windows before and after the adjustment corresponding to the same behavior change point are compared side by side. For each indicator, the difference value is calculated by subtracting the statistical value of the previous window from the statistical value of the later window, and then filled into the before-and-after difference comparison table.
[0083] Step 303: Based on statistical testing methods, perform statistical significance testing on the performance indicator data to obtain the significance test results.
[0084] Specifically, statistical testing methods refer to mathematical statistical methods used to infer whether differences observed in sample data also exist in the population, rather than being caused by random fluctuations. Common methods include comparing the differences between the means of two groups of data and nonparametric tests, used to determine whether the differences are statistically significant. The significance test result refers to the conclusive output obtained after applying statistical testing methods, including the value of the test statistic and a crucial p-value. The p-value represents the probability that the observed difference is purely due to randomness, and also includes a binary judgment conclusion, optionally indicating whether the difference is significant or not significant. For each point of behavioral change and its corresponding pre-adjustment and post-adjustment performance index data, the terminal performs the following operations for each key performance index: All data points of the index within the pre-adjustment window are considered sample A, and all data points of the index within the post-adjustment window are considered sample B. Based on the data distribution characteristics, an appropriate statistical test method is selected to perform hypothesis testing on samples A and B. The null hypothesis is typically set as statistically indistinguishable between the performance indices of the two windows. After the test, the test statistic and p-value are output. If the p-value is less than the set significance level, the null hypothesis is rejected, and the performance difference before and after adjustment is considered statistically significant.
[0085] Step 304: Based on the before-and-after difference comparison table and the significance test results, map the adjustment effect of the behavior change points to the corresponding levels to obtain the effect direction judgment table.
[0086] The "level" refers to the category label used for the final qualitative classification of the strategy adjustment effect. It is an ordered discrete set, which can be exemplified as: significant positive effect, slight positive effect, no significant effect, slight negative effect, and significant negative effect. Each level has a clear definition standard, combining the direction of difference and statistical significance. The "Effect Direction Judgment Table" refers to a final classification result list. For each behavioral change point evaluated, it provides a clear, rule-based effect level label. This table concisely summarizes the final qualitative conclusion of each strategy adjustment event. The terminal, based on a predefined mapping rule, combines the quantitative differences in the before-and-after difference comparison table with the statistical conclusions in the significance test results to assign an effect level to each behavioral change point. For example, for a behavioral change point and a key performance indicator, the significance test results are first checked. If the p-value is >= 0.05, it is judged as having no significant effect regardless of the size of the difference. If the p-value is < 0.05, the difference value of the indicator in the before-and-after difference comparison table is checked. If the difference value is positive and exceeds a certain positive threshold, it is judged as a significant positive effect; if it is positive but below the threshold, it is judged as a slight positive effect. The same applies to negative effects. By combining the judgment results of multiple key indicators, the overall effect level of the behavioral change point is given, completing the final transformation from data to conclusion. Complex statistical data and test results are condensed into a clear, easy-to-understand, and easy-to-use qualitative judgment.
[0087] Step 305: Based on the effect direction judgment table, generate an evaluation of the effect of strategy adjustment.
[0088] The strategy adjustment effect evaluation refers to a structured, summary evaluation document or data summary, which includes the classification results in the effect direction judgment table, along with relevant summary statistics and descriptions. For example, the report might show the distribution ratio of various effect levels across all behavioral change points, the adjustment effect patterns for different types of environmental changes, and the overall percentage of effective adjustments. It provides a holistic description of the effects of all strategy adjustment events for the agent. The terminal uses the effect direction judgment table as the primary input to count and statistically analyze the effect levels of all behavioral change points, calculate the proportion of each effect level, and perform cross-analysis based on the type of environmental change event. For example, it might analyze that when facing changes in resource constraints, the agent's adjustment effect is mostly significantly positive, while when facing sudden changes in system state, the effect is mostly insignificant. The analysis results are then organized in a structured format to form the final strategy adjustment effect evaluation document or data summary.
[0089] This embodiment provides a panoramic view and quantitative summary of the effectiveness of agent policy adjustments, gives case-by-case conclusions for each adjustment, and reveals the overall trend and pattern of the agent's effectiveness in responding to change.
[0090] In one embodiment, based on the policy adjustment effect evaluation and behavior adjustment delay measurement, the overall adaptability of the agent under evaluation is assessed to obtain a dynamic adaptation analysis report, including:
[0091] Step 401: Based on the behavioral change points, merge the data of the strategy adjustment effect evaluation and the behavioral adjustment delay measurement to obtain the comprehensive strategy adjustment evaluation table.
[0092] The strategy adjustment comprehensive evaluation table is a structured data table with each behavior change point as an independent entry. It integrates multiple evaluation information about the same strategy adjustment event into a single record. Each record contains the timestamp of the behavior change point, the corresponding environmental change event type, the delay time of this adjustment, the final effect level of this adjustment, and other contextual information that may be inherited from the original data. It serves as the foundational dataset for subsequent overall statistics and pattern analysis. The terminal uses a list of uniquely identified behavior change points as the primary key or index. For each behavior change point in the list, it searches and extracts the effect level assigned to that behavior change point from the strategy adjustment effect evaluation results, and searches and extracts the corresponding delay time from the behavior adjustment delay measurement value list. It then merges the behavior change point, timestamp, effect level, delay time, and any associated environmental change event information that may be obtained from the response pairing table to form a complete comprehensive evaluation record. By traversing all behavior change points and summarizing all records, the strategy adjustment comprehensive evaluation table is generated.
[0093] Step 402: Based on the comprehensive evaluation table of strategy adjustment, calculate the average adjustment latency and adjustment success rate using the following formulas to obtain a summary of adaptive performance indicators:
[0094]
[0095]
[0096] in, To represent the average adjustment delay, n is the total number of behavior change points, and i is the index of the behavior change point. The adjustment delay time is the time for the i-th behavior change point. To adjust the success rate, Classify the adjustment effect at the i-th behavior change point. This is an indicator function.
[0097] Specifically, the average adjustment delay refers to the arithmetic mean of the adjustment delay times of all records in the policy adjustment comprehensive evaluation table, summarizing the average speed at which the agent responds to environmental changes using a single value. The adjustment success rate refers to the proportion of all policy adjustment events evaluated as effective or positive. The indicator function returns 1 for positive effects (e.g., significantly positive, slightly positive) and 0 for other categories; the success rate is the average of the indicator function values that return 1. The summary of adaptation performance metrics is a list of calculation results including core comprehensive metrics such as average adjustment delay and adjustment success rate. It is a highly summarized performance snapshot used to quickly measure and compare the response speed and effectiveness of the agent's dynamic adaptation. The terminal takes the comprehensive evaluation table of strategy adjustment as input, calculates the average adjustment delay, reads the delay time field of all records in the table, sums them and divides by the total number of records to calculate the adjustment success rate. It iterates through each record and applies an indicator function to judge the effect level field of each record. For example, if the effect level is significant positive effect or slight positive effect, the function value is 1; if it is no significant effect or negative effect, the function value is 0. The indicator function values of all records are added together and then divided by the total number of records to obtain the success rate. The two calculation results, together with other possible derived indicators, constitute the summary of adaptive performance indicators.
[0098] Step 403: Based on the comprehensive evaluation table of strategy adjustment, the strategy adjustment events are grouped to obtain the clustering results of adaptive behavior patterns.
[0099] Specifically, adaptive behavior patterns refer to the strategy adjustment methods exhibited by an agent in response to external changes, sharing certain common characteristics. These patterns are categories abstracted and generalized from multiple specific adjustment events. Adaptive behavior pattern clustering results refer to the grouping conclusions obtained by performing unsupervised clustering analysis or rule-based supervised grouping on the records in the comprehensive evaluation table of strategy adjustments, classifying all events into several clusters or categories with similar internal characteristics. The terminal performs exploratory pattern discovery, using the comprehensive evaluation table of strategy adjustments as input. Each record is treated as a multi-dimensional data point, with dimensions including latency, effect level, associated environmental event type, and pre-adjustment effectiveness level. Clustering algorithms or predefined business rules are applied to group these data points. The purpose of grouping is to make events within the same group as similar in characteristics as possible, while events in different groups differ significantly. Analyzing the inherent commonalities of these groups allows for naming each group and defining an adaptive behavior pattern.
[0100] Step 404: Based on the summary of adaptive performance indicators and the clustering results of adaptive behavior patterns, summarize the overall adaptive capability of the agent to be evaluated and obtain a dynamic adaptation analysis report.
[0101] Overall adaptability refers to the final qualitative conclusion, based on all the aforementioned analyses, regarding the agent's comprehensive ability to maintain or improve its coordinated performance in a dynamically changing environment. It is a general answer to the question of whether the agent is good at adapting to change. The dynamic adaptation analysis report is a comprehensive and conclusive document containing core data from the summary of adaptation performance indicators, integrating deep-seated behavioral characteristics revealed by the clustering results of adaptation behavior patterns, and providing a structured evaluation of the agent's overall adaptability. The report typically includes a comprehensive performance score, main strengths and weaknesses, typical behavioral patterns, differences in adaptation performance in different scenarios, and suggestions for improvement. The summary of terminal adaptation performance indicators provides a macro-level performance assessment, including whether the average adjustment latency is short or long, and whether the adjustment success rate is high or low. In-depth analysis is then conducted using clustering results of adaptive behavior patterns to determine the uniformity of the agent's performance and the specific patterns or scenarios where successes or failures are concentrated. For example, although the average success rate is acceptable, clustering results show that all failures are concentrated on events involving sudden changes in system state—a key weakness. By combining quantitative indicators with pattern insights, a comprehensive overview of the agent's adaptive capabilities is summarized in text, highlighting its strengths, weaknesses, and unique behavioral characteristics, ultimately forming a dynamic adaptation analysis report.
[0102] This embodiment reveals the intrinsic structure and heterogeneity of the adaptive behavior of intelligent agents, and transforms a series of data analysis results into a comprehensive evaluation report, providing the final and authoritative basis for a comprehensive understanding of the dynamic adaptive capabilities of intelligent agents.
[0103] In one embodiment, before generating a comprehensive coordination evaluation report based on the dynamic adaptation analysis report and the coordination effectiveness numerical table, the method further includes:
[0104] Step 501: Traverse the coordination efficiency value table, filter out the timestamps corresponding to the coordination efficiency indicators that are lower than the predefined industry benchmark values, and obtain the list of inefficient times.
[0105] The predefined industry benchmark value refers to the minimum threshold set for each coordination effectiveness indicator, representing the acceptable or average level within that field. It serves as a reference standard, typically based on historical data, theoretical optimal values, or recognized industry standards, and is used to determine whether the agent's performance meets the benchmark. The inefficient time list is an ordered list containing specific time points. These time points correspond to the moments when the actual measured value of one or more effectiveness indicators in the coordination effectiveness value table is lower than their corresponding industry benchmark value. Each entry in the list usually includes a specific timestamp and the name of the indicator that failed to meet the benchmark at that moment. The terminal reads the coordination effectiveness value table, which records the specific values of each effectiveness indicator at different time points. It checks each pre-selected effectiveness indicator used for benchmark comparison row by row. For each indicator, its actual value at that time point is compared with the pre-set industry benchmark value. If the actual value of any indicator is lower than its corresponding benchmark value, the timestamp corresponding to that row of data is recorded and output. After traversing the entire value table, all recorded timestamps constitute the inefficient time list.
[0106] Step 502: Based on the list of inefficient times, extract the environmental feature subset corresponding to each timestamp from the coordinated behavior feature dataset to obtain the inefficient environmental features.
[0107] The Coordination Behavior Feature Dataset refers to the dataset generated in the preliminary evaluation steps, containing various raw features of the agent during task execution. It includes features describing the agent's own behavior and features describing the task environment state, such as the task difficulty level, available resources, the state of other agents, and communication quality. Inefficient Environment Features refer to all feature data extracted from the Coordination Behavior Feature Dataset specifically for each timestamp in the Inefficient Time List, describing the environmental state at that moment. It is a collection of data fragments, each fragment depicting the specific environmental conditions at the time of an inefficient performance. The terminal uses the Inefficient Time List as a query index, processing each timestamp in the list sequentially. For each timestamp, it searches the Coordination Behavior Feature Dataset to locate the data record row corresponding to that precise time point. From that row, it filters out all feature fields describing the environmental state, such as current task complexity, system load rate, and available bandwidth. The extracted feature data constitutes the environmental feature subset corresponding to that inefficient event. The environmental feature subsets extracted from all timestamps are then summarized to obtain the Inefficient Environment Feature Set.
[0108] Step 503: Input the inefficient environment features into the predefined expert rule base to obtain the standard response strategy, and compare the differences between the agent coordination strategy corresponding to the inefficient environment features and the standard response strategy to obtain the behavior difference analysis matrix.
[0109] The expert rule base refers to a knowledge base that encapsulates domain knowledge and best practices, containing a series of rules in the form of "if...then..." or a computable decision model. Its input is environmental state characteristics, and its output is a coordination strategy suggestion considered optimal or standard in that environment, representing the ideal or proven effective decision logic for that task scenario. The standard response strategy refers to the ideal coordination strategy that should be adopted for that specific environment, derived from the expert rule base after querying it and taking inefficient environmental characteristics as input. It represents the theoretically appropriate action plan that an expert-level agent would take under the same adverse environment. The agent coordination strategy refers to the coordination strategy inferred in the preliminary evaluation steps and actually adopted by the agent being evaluated at the moments marked in the inefficient time list during actual operation. The behavior difference analysis matrix is a structured table used to systematically compare the differences between the standard response strategy and the agent coordination strategy. The rows of the matrix may represent different strategy dimensions, and the columns represent different instances of inefficient events. Each cell records a specific description of the differences between the agent's actual strategy and the standard strategy in that strategy dimension. For each subset of environmental features in the inefficient environment feature set, the terminal takes it as input parameter and calls a predefined expert rule base. The rule base outputs a recommended strategy, i.e. a standard response strategy, based on its built-in logic. For the same inefficient event, the terminal obtains the agent coordination strategy actually adopted at that time. On multiple preset and comparable strategy dimensions, the standard strategy and the actual strategy are compared item by item, and the differences found in the comparison are recorded in a matrix.
[0110] Step 504: Based on behavioral difference analysis, analyze the reasoning process of the agent to be evaluated when making decisions, and obtain the decision logic defects.
[0111] Specifically, the decision-making reasoning process refers to the internal computational, judgmental, and selection chain that an intelligent agent undergoes from perceiving the environment to forming the final coordinated strategy. This involves the internal working mechanisms of its decision-making model, algorithm, rules, or neural network. Decision logic defects refer to fundamental errors, deficiencies, or limitations in the agent's decision-making reasoning process, and are the intrinsic causes of systematic behavioral biases. The terminal uses a behavioral difference analysis matrix as its primary input, but the analysis delves deeper into the internal reasoning logic beyond surface behavioral differences. The analysis combines the agent's specific implementation model to retrospectively analyze the differences recorded in the matrix. For example, given the difference between a standard policy suggestion broadcast request and an actual policy point-to-point request, the analysis needs to search why the agent's decision logic made a point-to-point choice. Is it because the communication rule base lacks broadcast entries, the value function incorrectly assesses the cost of broadcast, or the opponent model is inaccurate, misjudging point-to-point as more effective? By systematically searching all differences, the fundamental errors or omissions at the agent's decision-making model or algorithm level that lead to these common differences can be inferred.
[0112] Step 505: Based on decision logic analysis, a coordination capability bottleneck report is obtained; the coordination capability bottleneck report is used to assist in generating a comprehensive coordination assessment report.
[0113] Specifically, the Coordination Capability Bottleneck Report is a summary document that systematically summarizes and elaborates on the key obstacles and core weaknesses identified in the decision logic defect analysis that restrict the improvement of the coordination capability of the agent to be evaluated. It clearly lists the most significant bottlenecks and links them to the specific decision logic defects that lead to these bottlenecks. The terminal takes the list of decision logic defects as input, summarizes, classifies, and prioritizes these defects, merges related and similar defects, and extracts higher-level capability shortcomings. Optionally, it summarizes the two defects of ignoring resource competition risks and not considering load balancing in task allocation as a lack of resource collaborative optimization capability. It assesses the severity and prevalence of the impact of capability shortcomings on overall coordination efficiency, identifies the most critical bottlenecks, and writes a structured report that clearly states what these bottlenecks are and how they lead to inefficiency.
[0114] This embodiment attributes the agent's inefficient performance to the deviation between its specific decision-making behavior and its ideal decision-making behavior, revealing the deep-seated and systemic reasons for the agent's inefficient behavior, and pointing out the key obstacles to improving the agent's capabilities. This information serves as an important input to assist in generating a comprehensive coordination evaluation report.
[0115] In one embodiment, the differences between the agent coordination strategy corresponding to the inefficient environment characteristics and the standard response strategy are compared to obtain a behavior difference analysis matrix, including:
[0116] Step 601: Match and align the agent coordination strategy and the standard response strategy to obtain an alignment pairing table.
[0117] The agent coordination strategy is an inference from the agent's actual behavior records during the preliminary evaluation steps, representing the sequence of inherent decision-making rules or behavioral patterns actually adopted at various time points or in different situations. The standard response strategy is a theoretically ideal or optimal decision-making rule generated by inputting inefficient environment features into an expert rule base, for each specific inefficient environment scenario. The alignment pairing table is a structured mapping table whose core function is to establish a one-to-one correspondence between agent coordination strategies and standard response strategies under the same or comparable conditions. Each row in the table represents a pairing, including: a pairing identifier, a description of the corresponding environmental features, the agent's actual coordination strategy in that situation, and the standard response strategy recommended by the expert system for that situation. The terminal performs a context matching process to ensure that the comparison is conducted under the same conditions. It traverses all the scenarios that need to be analyzed. For each scenario, it finds the policy record actually adopted by the agent in the inferred agent coordination policy sequence and finds the standard policy record that outputs for the exact same environmental feature input from the generated standard response policy set. The two policy records are paired and recorded in a table. The key to matching is to ensure that the environmental state and task goal addressed by the two policies are completely consistent, so as to conduct a fair and meaningful difference comparison.
[0118] Step 602: Based on the alignment pairing table, compare the differences between the agent's coordination strategy and the standard response strategy to obtain a list of decision difference points.
[0119] Specifically, the decision difference list refers to a detailed record of the differences between each pair of agent coordination strategies and the standard response strategy in specific decision dimensions. Each item in the list is a specific difference description, including the pair to which the difference belongs, the specific strategy dimension, the agent's actual decision value, the recommended decision value of the standard strategy, and a brief description of the difference. The terminal performs a fine-grained comparison item by item, reading each pairing record in the alignment pairing table. For each pair of strategies, it performs a comparison item by item on preset, decomposable strategy comparison dimensions. The dimensions are predefined and used to deconstruct a complete coordination strategy. For example, dimensions may include the selection of collaborating objects, the format of message content, the timing of action execution, the priority of resource usage, etc. On each dimension, the parameters or values of that dimension in the agent's strategy are compared with the parameters or values of the corresponding dimension in the standard strategy. If an inconsistency is found, a decision difference record is generated, clearly indicating which dimension is different and how. After traversing all pairs and all preset dimensions, all differences are summarized to form the decision difference list.
[0120] Step 603: Based on the list of decision discrepancies, calculate the overall discrepancy index to obtain a summary of the overall discrepancy index; the overall discrepancy index includes at least one of decision consistency rate, response time deviation, and key decision error rate.
[0121] Specifically, the overall difference index refers to a statistical measure used to quantify the degree of difference between an agent's strategy and a standard strategy at a macro level. It is a comprehensive value that summarizes the breadth, frequency, or importance of the difference. Key metrics include: Decision Consistency Rate (DCPR), which measures the proportion of agent decisions that align with standard decisions across all compared decision points, indicating the degree of fit between overall behavior and the ideal pattern; Response Time Deviation, which calculates the average deviation or root mean square error between the agent's actual time parameters and standard time parameters when the strategy involves time parameters, measuring the difference in timeliness; and Critical Decision Error Rate, which measures the proportion of times the agent makes decisions contrary to standard decisions on predefined critical decision dimensions that decisively influence task success or failure, out of the total number of critical decisions, measuring the frequency of errors on critical issues. The overall difference index summary refers to the collection of specific values obtained after calculating the above-mentioned overall difference indices. The terminal performs calculations based on a list of decision discrepancies and an alignment pairing table. For example, for the decision consistency rate, the total number of discrepancies in the list is recorded as the total number of discrepancies, and the total number of comparison points equals the number of alignment pairs multiplied by the preset number of strategy comparison dimensions. For response time deviation, all discrepancies involving time parameters are selected from the decision discrepancy list, and the difference between the actual value and the standard value for each discrepancy point is calculated. The average or root mean square of these differences is then calculated. For the critical decision error rate, which decision dimensions are considered critical decisions are clearly defined. The comparison results of all relevant dimensions are iterated, and the number of times the agent's choice on that dimension is inconsistent with the standard choice is counted. This number is then divided by the total number of critical decision comparisons. These calculation results are then compiled to obtain a summary of the overall discrepancy index, achieving a leap from specific discrepancies to macroscopic discrepancies. The effect is to condense a large number of scattered, specific discrepancies into several numerical indicators with global significance that are easy to understand and compare.
[0122] Step 604: Based on the summary of overall difference indicators, obtain the behavioral difference analysis matrix.
[0123] The behavioral difference analysis matrix is a structured data table or document that integrates all the core results of difference analysis. It is a comprehensive analytical view and typically includes at least the basic information of alignment and pairing; a summary of the main differences extracted from the list of decision difference points and categorized by context or dimension; and key quantitative indicators derived from the overall difference index summary. Rows represent different strategy dimensions or different inefficient events, and columns include actual strategies, standard strategies, specific difference descriptions, and overall indicators. The terminal uses the alignment and pairing table as an index and framework, filling in the detailed difference descriptions corresponding to the decision difference point list in a structured manner. The calculated overall difference index summary is placed prominently in the matrix as summary information. All this information is organized into a logically clear and easily searchable composite data structure or document to obtain the behavioral difference analysis matrix.
[0124] This embodiment preserves the details of specific differences, facilitating in-depth tracing; it also provides a global quantitative assessment, making it easier for high-level judgment, thereby improving the accuracy and reliability of the assessment.
[0125] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0126] Based on the same inventive concept, this application also provides a power large-scale intelligent agent operation evaluation device for implementing the power large-scale intelligent agent operation evaluation method described above. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more power large-scale intelligent agent operation evaluation device embodiments provided below can be found in the limitations of the power large-scale intelligent agent operation evaluation method described above, and will not be repeated here.
[0127] In one exemplary embodiment, such as Figure 2 As shown, a power large-scale model intelligent agent operation evaluation device 700 is provided, comprising:
[0128] The recording module 701 is used to generate a task flow based on the evaluation purpose, and input the task flow into the agent to be evaluated to obtain the agent's behavior record;
[0129] The feature module 702 is used to extract features from the agent's behavior records to obtain a coordination behavior feature dataset, and based on the coordination behavior feature dataset, to infer the coordination strategy of the agent to be evaluated and obtain the agent's coordination strategy.
[0130] The quantification module 703 is used to quantify the indicators of coordination effectiveness based on the coordination behavior feature dataset and obtain a coordination effectiveness numerical table.
[0131] Dynamic module 704 is used to analyze the dynamic adaptability of the agent to be evaluated based on the agent coordination strategy and coordination effectiveness numerical table, and obtain a dynamic adaptability analysis report.
[0132] Report module 705 is used to generate a comprehensive coordination assessment report based on the dynamic adaptation analysis report and the coordination effectiveness numerical table; the comprehensive coordination assessment report is used to assist in assessing the current capabilities of the agent to be assessed.
[0133] Furthermore, dynamic module 704 is also used for:
[0134] Environmental change events are extracted from the task flow, and based on the agent coordination strategy, the time points when the agent to be evaluated adjusts its strategy are identified to obtain the behavior change points; the environmental change events include at least one of the following: changes in task attributes, changes in system state, and changes in resource constraints.
[0135] By matching environmental change events and behavioral change points in time, a response matching table is obtained;
[0136] Based on the response pairing table, the time difference between each environmental change event and behavioral change point is calculated to obtain the behavioral adjustment delay measurement value.
[0137] Based on the coordination effectiveness numerical table and behavioral change points, the dynamic response effect of strategy adjustment events is judged, and the strategy adjustment effect evaluation is obtained.
[0138] Based on the evaluation of the effect of policy adjustment and the measurement of the delay in behavior adjustment, the overall adaptability of the agent under evaluation is assessed, and a dynamic adaptation analysis report is obtained.
[0139] Furthermore, dynamic module 704 is also used for:
[0140] Define analysis windows of fixed length before and after the behavior change point, and extract all coordination effectiveness indicators within the analysis window based on the coordination effectiveness numerical table to obtain effectiveness indicator data.
[0141] Calculate the statistical measures of the performance indicators to obtain the performance indicator statistical measures table; based on the points of behavioral change, compare and analyze the statistical measures of the indicators between the windows to obtain the before-and-after difference comparison table.
[0142] Based on statistical testing methods, statistical significance tests were performed on the performance indicator data to obtain the significance test results;
[0143] Based on the before-and-after difference comparison table and the significance test results, the adjustment effect of the behavioral change points is mapped to the corresponding level to obtain the effect direction judgment table.
[0144] Based on the effect direction judgment table, an evaluation of the effect of strategy adjustment is generated.
[0145] Furthermore, dynamic module 704 is also used for:
[0146] Based on the points of behavioral change, the data of the strategy adjustment effect evaluation and the behavioral adjustment delay measurement are combined to obtain a comprehensive strategy adjustment evaluation table.
[0147] Based on the comprehensive evaluation table of strategy adjustments, the average adjustment latency and adjustment success rate are calculated using the following formulas to obtain a summary of adaptive performance indicators:
[0148]
[0149]
[0150] in, To represent the average adjustment delay, n is the total number of behavior change points, and i is the index of the behavior change point. The adjustment delay time is the time for the i-th behavior change point. To adjust the success rate, Classify the adjustment effect at the i-th behavior change point. For indicator functions;
[0151] Based on the comprehensive evaluation table of strategy adjustment, strategy adjustment events are grouped to obtain clustering results of adaptive behavior patterns.
[0152] Based on the summary of adaptive performance indicators and the clustering results of adaptive behavior patterns, the overall adaptive capability of the agent to be evaluated is summarized, and a dynamic adaptive analysis report is obtained.
[0153] Furthermore, the device also includes a bottleneck module for:
[0154] Traverse the coordination efficiency value table, filter out the timestamps corresponding to the indicators of coordination efficiency that are lower than the predefined industry benchmark value, and obtain the list of inefficient times;
[0155] Based on the list of inefficient times, a subset of environmental features corresponding to each timestamp is extracted from the coordinated behavior feature dataset to obtain inefficient environmental features.
[0156] Inefficient environment characteristics are input into a predefined expert rule base to obtain standard response strategies. The differences between the agent coordination strategies corresponding to inefficient environment characteristics and the standard response strategies are compared to obtain a behavior difference analysis matrix.
[0157] Based on behavioral difference analysis, the reasoning process of the agent under evaluation during decision-making is analyzed to identify decision logic defects.
[0158] Based on decision-making logic analysis, a coordination capacity bottleneck report is obtained; the coordination capacity bottleneck report is used to assist in generating a comprehensive coordination assessment report.
[0159] Furthermore, the bottleneck module is also used for:
[0160] The agent coordination strategy and the standard response strategy are matched and aligned to obtain an alignment pairing table;
[0161] Based on the alignment and pairing table, the differences between the agent coordination strategy and the standard response strategy are compared to obtain a list of decision difference points;
[0162] Based on the list of decision-making discrepancies, calculate the overall discrepancy index to obtain a summary of the overall discrepancy index; the overall discrepancy index includes at least one of the following: decision consistency rate, response time deviation, and critical decision error rate.
[0163] Based on the summary of overall difference indicators, a behavioral difference analysis matrix is obtained.
[0164] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above-described power large-scale model intelligent agent operation evaluation method.
[0165] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0166] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0167] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A method for evaluating the operation of a large-scale power model intelligent agent, characterized in that, The method includes: Based on the evaluation purpose, a task flow is generated and input into the agent to be evaluated to obtain the agent's behavior record; Feature extraction is performed on the agent's behavior records to obtain a coordination behavior feature dataset. Based on the coordination behavior feature dataset, the coordination strategy of the agent to be evaluated is inferred to obtain the agent's coordination strategy. Based on the aforementioned coordination behavior feature dataset, indicators of coordination effectiveness are quantified to obtain a coordination effectiveness numerical table. Based on the agent coordination strategy and the coordination performance table, the dynamic adaptability of the agent to be evaluated is analyzed to obtain a dynamic adaptability analysis report. Based on the dynamic adaptation analysis report and the coordination effectiveness numerical table, a comprehensive coordination evaluation report is generated; the comprehensive coordination evaluation report is used to assist in evaluating the current capabilities of the agent to be evaluated.
2. The method according to claim 1, characterized in that, The dynamic adaptability of the agent to be evaluated is analyzed based on the agent coordination strategy and the coordination effectiveness numerical table to obtain a dynamic adaptability analysis report, including: Environmental change events are extracted from the task flow, and based on the agent coordination strategy, the time points of the policy adjustment events of the agent to be evaluated are identified to obtain the behavior change points; the environmental change events include at least one of task attribute changes, system state changes, and resource constraint changes; The environmental change events and behavioral change points are matched over time to obtain a response pairing table; Based on the response pairing table, the time difference between each environmental change event and the behavioral change point is calculated to obtain a behavioral adjustment delay measurement value. Based on the coordination effectiveness numerical table and the behavioral change points, the dynamic response effect of the strategy adjustment event is judged to obtain the strategy adjustment effect evaluation. Based on the evaluation of the policy adjustment effect and the measurement of the behavior adjustment delay, the overall adaptability of the agent to be evaluated is assessed, and the dynamic adaptation analysis report is obtained.
3. The method according to claim 2, characterized in that, The step of determining the dynamic response effect of the strategy adjustment event based on the coordination effectiveness numerical table and the behavioral change points, and obtaining a strategy adjustment effect evaluation, includes: Define analysis windows of fixed length before and after the behavior change point, and extract all the coordination effectiveness indicators within the analysis windows based on the coordination effectiveness value table to obtain effectiveness indicator data. Calculate the performance indicator statistics for the performance indicator data to obtain a performance indicator statistics table; compare the performance indicator statistics between the analysis windows based on the behavioral change points to obtain a before-and-after difference comparison table. Based on statistical testing methods, the performance index data were subjected to statistical significance testing to obtain the significance test results; Based on the before-and-after difference comparison table and the significance test results, the adjustment effect of the behavior change point is mapped to the corresponding level to obtain the effect direction judgment table; Based on the effect direction judgment table, an evaluation of the effect of the strategy adjustment is generated.
4. The method according to claim 2, characterized in that, The overall adaptability of the agent under evaluation is assessed based on the policy adjustment effect evaluation and the behavior adjustment delay measurement value, resulting in the dynamic adaptation analysis report, including: Based on the behavioral change points, the data of the strategy adjustment effect evaluation and the behavioral adjustment delay measurement are combined to obtain a comprehensive strategy adjustment evaluation table; Based on the aforementioned strategy adjustment comprehensive evaluation table, the average adjustment latency and adjustment success rate are calculated using the following formulas to obtain a summary of adaptive performance indicators: in, To represent the average adjustment delay, n is the total number of behavior change points, and i is the index of the behavior change point. The adjustment delay time is the time for the i-th behavior change point. To adjust the success rate, Classify the adjustment effect at the i-th behavior change point. For indicator functions; Based on the comprehensive evaluation table for strategy adjustment, the strategy adjustment events are grouped to obtain clustering results of adaptive behavior patterns. Based on the summary of the adaptive performance indicators and the clustering results of the adaptive behavior patterns, the overall adaptive capability of the agent to be evaluated is summarized, and the dynamic adaptation analysis report is obtained.
5. The method according to claim 1, characterized in that, Before generating the comprehensive coordination evaluation report based on the dynamic adaptation analysis report and the coordination effectiveness numerical table, the process also includes: Traverse the coordination efficiency value table, filter out the timestamps corresponding to the indicators of coordination efficiency that are lower than the predefined industry benchmark value, and obtain the list of inefficient times; Based on the list of inefficient times, a subset of environmental features corresponding to each timestamp is extracted from the coordinated behavior feature dataset to obtain inefficient environmental features. The inefficient environment characteristics are input into a predefined expert rule base to obtain a standard response strategy. The differences between the agent coordination strategy corresponding to the inefficient environment characteristics and the standard response strategy are compared to obtain a behavior difference analysis matrix. Based on the behavioral difference analysis, the reasoning process of the agent to be evaluated during decision-making is analyzed to identify decision logic defects. Based on the decision logic analysis, a coordination capability bottleneck report is obtained; the coordination capability bottleneck report is used to assist in generating the comprehensive coordination assessment report.
6. The method according to claim 5, characterized in that, The comparison of the agent coordination strategy and the standard response strategy corresponding to the inefficient environmental characteristics yields a behavioral difference analysis matrix, including: The agent coordination strategy and the standard response strategy are matched and aligned to obtain an alignment pairing table; Based on the alignment pairing table, the differences between the agent coordination strategy and the standard response strategy are compared to obtain a list of decision difference points; Based on the list of decision discrepancies, an overall discrepancy index is calculated to obtain a summary of the overall discrepancy index; the overall discrepancy index includes at least one of decision consistency rate, response time deviation, and key decision error rate. Based on the summary of the overall difference indicators, the behavioral difference analysis matrix is obtained.
7. A power large-scale model intelligent agent operation evaluation device, characterized in that, The device includes: The recording module is used to generate a task flow based on the evaluation purpose, and input the task flow into the agent to be evaluated to obtain the agent's behavior record; The feature module is used to extract features from the agent's behavior records to obtain a coordination behavior feature dataset, and based on the coordination behavior feature dataset, to infer the coordination strategy of the agent to be evaluated, and obtain the agent coordination strategy. The quantification module is used to quantify the indicators of coordination effectiveness based on the coordination behavior feature dataset, and obtain a coordination effectiveness numerical table. The dynamic module is used to analyze the dynamic adaptability of the agent to be evaluated based on the agent coordination strategy and the coordination performance value table, and to obtain a dynamic adaptability analysis report. The reporting module is used to generate a comprehensive coordination evaluation report based on the dynamic adaptation analysis report and the coordination effectiveness numerical table; the comprehensive coordination evaluation report is used to assist in evaluating the current capabilities of the agent to be evaluated.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.