Interpretable evaluation method and system for the effectiveness of multi-agent reinforcement learning algorithms based on Shapley additive interpretation

By constructing a hierarchical evaluation index system and an improved Deep-SHAP method, the interpretability problem of multi-agent reinforcement learning algorithms in practical applications is solved, high-precision evaluation and feature importance ranking are achieved, and the interpretability and evaluation accuracy of the algorithm are improved.

CN120524843BActive Publication Date: 2025-09-30NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511031628.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-30
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning algorithms lack interpretability in practical application scenarios. Traditional evaluation methods cannot effectively quantify the dynamics and behavior of target vessels, and the assumption of feature independence is not applicable to actual situations.

Method used

A hierarchical evaluation index system is constructed based on the Shapley additive interpretation method. Combined with the multi-layer perceptron model and the improved loss function, the empirical conditional distribution is used to model the dependency between features. The Shapley value is calculated by the Deep-SHAP method for interpretable evaluation.

Benefits of technology

It achieves scientific evaluation of multi-agent reinforcement learning algorithms in different application scenarios, improves the accuracy and precision of the evaluation, provides explainability of algorithm behavior, and obtains accurate ranking of feature importance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524843B_ABST
    Figure CN120524843B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for interpretable evaluation of the performance of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation. The method comprises: establishing a simulation environment for a multi-agent reinforcement learning algorithm at sea, establishing an evaluation index system adapted to the application scenario of the multi-agent reinforcement learning algorithm; using a multi-layer perceptron to synthesize the various indicators in the evaluation index system to obtain an algorithm performance evaluation prediction value, and improving the loss function design in the algorithm performance evaluation learning and training process; using the Deep‑SHAP method improved based on empirical conditional distribution, and performing interpretable analysis on the performance evaluation prediction value of the multi-layer perceptron based on the feature attribution idea of ​​the Shapley method, while visualizing the feature importance ranking. The present invention not only takes into account the independence between indicators in actual application scenarios, but also integrates neural network models and objective data, and can quickly achieve interpretable evaluation of multi-agent reinforcement learning algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research field of interpretable performance of marine multi-agent reinforcement learning algorithms, and in particular to a method and system for interpretable evaluation of the performance of multi-agent reinforcement learning algorithms based on Shapley additive interpretation. Background Art

[0002] With the rapid development of the economy and unmanned control technology, the demand for unmanned boats in the civilian sector continues to grow. As a new type of surface vehicle, unmanned surface boats (UAVs) can cost-effectively carry out operations in vast ocean areas. As the advantages of UAVs in ocean exploration become increasingly prominent, swarm operations of UAVs will become a key development trend in future ocean exploration. At the same time, the demand for multi-agent reinforcement learning technology in the field of UAV applications is increasing, and multi-agent reinforcement learning algorithms for UAV collaborative operations are rapidly developing. Through information exchange and collaborative operations between multiple UAVs, overall maritime operational efficiency can be improved and operating costs reduced.

[0003] At present, research hotspots in the field of multi-agent reinforcement learning mostly tend to be on the design and optimization of multi-agent reinforcement learning algorithms, and often underestimate or ignore the interpretability of algorithm behavior. There is little research on how to obtain the interpretability of multi-agent reinforcement learning algorithms applied in practical scenarios, and there is no mature and complete algorithm interpretability evaluation process system yet.

[0004] Understanding the behavior of multi-agent reinforcement learning algorithms is crucial for their continuous improvement and optimization. The interpretability of multi-agent reinforcement learning algorithms in practical application scenarios faces the following difficulties: First, when constructing algorithm evaluation metrics, the dynamics and behavior of target vessels are rarely quantified, and the evaluation metrics are not constructed in a hierarchical manner, resulting in an incomplete and unclear system. Second, traditional interpretable methods for algorithm performance are image-based, interpreting pixels in visual images and applying them to multi-agent reinforcement learning algorithms, which is not flexible enough. Finally, the traditional Shapley interpretation method assumes independence between features, but in reality, features often have dependencies and influence each other. Therefore, the traditional Shapley interpretation method is not suitable for interpreting multi-agent reinforcement learning algorithms in practical application scenarios. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for interpretable evaluation of the effectiveness of multi-agent reinforcement learning algorithms based on Shapley additive interpretation.

[0006] The technical solution for achieving the purpose of the present invention is: a method for interpretable evaluation of the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation, comprising the following steps:

[0007] Step 1: Establish a simulation environment for a multi-agent reinforcement learning algorithm at sea and establish an evaluation index system that is suitable for the application scenario of the multi-agent reinforcement learning algorithm.

[0008] Step 2: Based on the maritime multi-agent reinforcement learning algorithm simulation environment and evaluation index system established in step 1, simulate the maritime multi-agent reinforcement learning algorithm simulation environment algorithm under different conditions in the maritime simulation environment, collect raw data of evaluation indicators, use the raw data of indicators to obtain algorithm evaluation values ​​based on expert knowledge, and form the collected raw data of indicators and algorithm evaluation values ​​into evaluation samples;

[0009] Step 3: Build an evaluation prediction model based on a multi-layer perceptron to imitate expert knowledge and predict the algorithm performance evaluation value; train the evaluation prediction model based on the evaluation samples obtained in step 2, and improve the design of the loss function of the evaluation prediction model learning and training to obtain a better learning and training model to predict the algorithm performance evaluation value of the sample to be evaluated;

[0010] Step 4: Based on the evaluation samples obtained in step 2, select data to form a background data set; model the dependencies between features based on the empirical conditional distribution, and form a conditional distribution that satisfies feature independence in the background data set;

[0011] Step 5: Based on the empirical conditional distribution obtained in step 4, the empirical conditional distribution and model gradient are introduced into the Deep-SHAP method to obtain the improved Deep-SHAP interpretable method based on the empirical conditional distribution, thereby calculating the Shapley value of the predicted value of the algorithm performance evaluation of all indicators in the evaluation index system;

[0012] Step 6: Visualize the feature importance ranking based on the Shapley value obtained in step 5.

[0013] A multi-agent reinforcement learning algorithm performance interpretable evaluation system based on Shapley additive interpretation is used to implement the above evaluation method. The system includes:

[0014] The first module is used to establish a marine algorithm simulation environment and evaluation index system that is suitable for the application scenarios of multi-agent reinforcement learning algorithms;

[0015] The second module, based on the algorithm simulation environment and evaluation index system established in the first module, simulates the multi-agent reinforcement learning algorithm in different actual scenarios in the marine algorithm simulation environment, collects the original data of the evaluation indicators, and uses the original data of the indicators to obtain the algorithm evaluation value based on expert knowledge. The collected original data of the indicators and the algorithm evaluation value constitute the evaluation sample;

[0016] The third module builds an evaluation prediction model based on a multi-layer perceptron, imitating expert knowledge to predict the algorithm performance evaluation value; the evaluation prediction model is trained based on the evaluation samples obtained in the second module, and the loss function design of the evaluation prediction model learning and training is improved to obtain a better learning and training model to predict the algorithm performance evaluation value of the samples to be evaluated;

[0017] The fourth module selects data to form a background data set based on the evaluation samples obtained in the second module, models the dependencies between features based on the empirical conditional distribution, and forms a conditional distribution that satisfies feature independence in the background data set;

[0018] The fifth module calculates the Shapley value of the algorithm performance evaluation prediction value of all indicators in the evaluation index system for the samples to be evaluated based on the background data set obtained in the fourth module and the improved Deep-SHAP method;

[0019] The sixth module visualizes the feature importance ranking based on the Shapley value obtained in the fifth module.

[0020] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for interpretable evaluation of the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation is implemented.

[0021] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned interpretable evaluation method for the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation.

[0022] Compared with the prior art, the present invention can achieve the following technical effects by adopting the above technical solutions: (1) The present invention comprehensively considers the principles of completeness, hierarchy and accuracy, and constructs a scientific and complete multi-agent reinforcement learning algorithm performance evaluation index system for multi-agent reinforcement learning algorithms in different practical application scenarios; (2) The proposed algorithm performance evaluation prediction model based on multi-layer perceptron makes full use of the real sample data in the actual simulation environment, and at the same time improves the design of the loss function in the training process, effectively improving the prediction accuracy; (3) The proposed Deep-SHAP method based on the improved empirical conditional distribution solves the problem that the traditional Shapley method assumes the independence of features in actual scenarios, and uses the empirical conditional distribution to model the dependence between features in actual situations, thereby calculating a more accurate Shapley value. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a multi-agent reinforcement learning intelligent algorithm performance evaluation index system based on offshore fishery supervision scenarios.

[0024] Figure 2 This is a principle block diagram of the Shapley additive interpretation-based multi-agent reinforcement learning algorithm performance explainable method of the present invention.

[0025] Figure 3 This is a structural diagram of a multi-layer perceptron prediction model used to predict the performance evaluation value of an algorithm.

[0026] Figure 4 This is a feature importance ranking diagram of a multi-agent reinforcement learning algorithm based on offshore fishery supervision scenarios. DETAILED DESCRIPTION

[0027] In response to the difficulty of interpreting the performance of multi-agent reinforcement learning algorithms at sea, the present invention proposes a method and system for interpretable evaluation of the performance of multi-agent reinforcement learning algorithms based on Shapley additive interpretation. By quantifying information such as the position and distance of the target ship, a complete algorithm evaluation index system applicable to different application scenarios is hierarchically constructed, overcoming the problems of incomplete index system and unclear structure faced in the algorithm performance evaluation process. At the same time, a multi-agent reinforcement learning algorithm performance evaluation prediction method based on multi-layer perceptron is proposed to improve the loss function of learning and training, obtain a high-precision prediction of the performance evaluation value of the multi-agent reinforcement learning algorithm to be evaluated, and provide a basis for the interpretability of subsequent algorithms. Finally, a Deep-SHAP interpretation method based on empirical conditional distribution is proposed to solve the problem of dependency between features in actual application scenarios, and can achieve high-precision interpretable prediction of multi-agent reinforcement learning algorithms. Based on the feature attribution idea of ​​the Shapley method, the feature importance ranking of the multi-agent reinforcement learning algorithm is obtained.

[0028] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0029] like Figure 1 As shown in the figure, the multi-agent reinforcement learning intelligent algorithm performance evaluation index system based on the offshore fishery supervision scenario mainly includes: a multi-agent reinforcement learning algorithm performance interpretable index system construction module, an evaluation index system indicator raw data collection module, an algorithm performance evaluation training module, a multi-agent reinforcement learning algorithm performance evaluation value prediction module to be evaluated, a loss function improvement design module, an empirical conditional distribution improved Depp-SHAP interpretable analysis module, and a feature importance ranking visualization module.

[0030] Combine Figure 2 The present invention proposes an interpretable evaluation method for the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation. The specific implementation steps are as follows:

[0031] Step 1: Establish a simulation environment for the marine multi-agent reinforcement learning algorithm and establish an evaluation index system that is suitable for the application scenarios of the multi-agent reinforcement learning algorithm.

[0032] The multi-agent reinforcement learning algorithm is applied to different practical scenarios, and a hierarchical evaluation index system is constructed from multiple perspectives. Considering the offshore fishery supervision scenario as an actual application scenario, a comprehensive evaluation of the actual performance of the multi-agent reinforcement learning algorithm in the offshore fishery supervision application scenario is conducted based on a hierarchical structure, combined with multiple factors such as task conditions, evaluation objects, and environmental complexity. From the application of the algorithm to the actual scenario, the operation execution efficiency, operation and safety capabilities, expulsion and interception capabilities, resource consumption, and algorithm performance in multiple dimensions, a hierarchical evaluation index system is constructed. The established hierarchical structure diagram is shown in the figure below. Figure 1 The detailed definitions of each indicator in the evaluation index system are as follows:

[0033] 1) Assignment completion time:

[0034]

[0035] In the process of offshore fishery supervision, when a patrol unmanned boat group is sailing, the time for the patrol boat group to start the task of driving away the target ship is , the effective interception time for the target ship is .

[0036] 2) Assignment completion rate

[0037]

[0038] In the process of offshore fishery supervision, the mission completion rate is defined as the interception rate of the target ship by the patrol unmanned boat cluster. It represents the total number of target ships. It indicates the number of target ships that the patrol unmanned boat successfully drove away / intercepted.

[0039] 3) Patrol boat group range

[0040]

[0041] in, It refers to the patrol unmanned boat in the process of offshore fishery supervision. 's driving range. The number of patrol unmanned boat clusters.

[0042] 4) Patrol boat group online rate

[0043]

[0044] In the process of offshore fishery supervision, the online rate of patrol boat groups is defined as the ratio of fault-free patrol unmanned boats in the patrol unmanned boat group after a dispersal / interception operation. It represents the total number of patrol unmanned boats. It indicates the number of patrol unmanned boats that successfully drove away / intercepted the target ship.

[0045] 5) Algorithm decision time

[0046]

[0047] Generally, the decision-making of the algorithm takes time. In offshore fishery supervision, the time for the algorithm to start using the established strategy is , the algorithm generates the decision position time of the first patrol unmanned boat as .

[0048] 6) Algorithm memory usage

[0049]

[0050] Indicates the maximum memory usage of the computing chip required for each decision control signal generated by the intelligent algorithm used by the patrol unmanned boat cluster during offshore fishery supervision. Represents the number of times the generation algorithm generates a decision.

[0051] 7) Total voyage of target ship

[0052]

[0053] During an operation where a patrol unmanned boat drives away / intercepts a target ship, multiple target ships may appear at the same time. Indicates the target vessel in the process of offshore fishery supervision 's driving range. is the number of all target ships in one operation mission.

[0054] 8) Illegal area approaches threshold

[0055]

[0056] in, Indicates the The closest distance between the target ship and the illegal area.

[0057] 9) Minimum safety distance of equipment

[0058]

[0059] In offshore fisheries supervision, Indicates one's patrol unmanned boat and The minimum distance to the target ship.

[0060] Step 2: Based on the maritime multi-agent reinforcement learning algorithm simulation environment and evaluation index system established in Step 1, simulate the maritime multi-agent reinforcement learning algorithm simulation environment algorithm under different conditions in the maritime simulation environment, collect the original data of the evaluation indicators, and use the original data of the indicators to obtain the algorithm evaluation value based on expert knowledge. The collected original data of the indicators and the algorithm evaluation value constitute the evaluation sample, which is specifically as follows:

[0061] Based on the constructed hierarchical evaluation index system and the marine multi-agent reinforcement learning algorithm simulation environment, the multi-agent reinforcement learning algorithms in different application scenarios are simulated by configuring the initial parameters in the marine simulation environment, and the original data of the evaluation indicators are collected. The upper and lower bounds of the indicator data are set according to the historical data of different algorithms. The weights are calculated according to the standardization of the indicator data, and the scores of each indicator are weighted to obtain the corresponding algorithm evaluation value. The collected original indicator data and algorithm evaluation values ​​of the evaluation samples constitute the evaluation samples.

[0062]

[0063] in, Indicates the There are 9 indicators in total, including operation completion time, operation completion rate, patrol boat group range, patrol boat group online rate, algorithm decision time, algorithm memory usage, target ship total range, illegal zone approach threshold, and equipment minimum safety distance; Indicates the The lower bound of the indicator, Indicates the The upper bound of the indicator, represents the algorithm performance evaluation value, represents the lower bound of the algorithm performance evaluation value, Indicates the upper bound of the algorithm performance evaluation value.

[0064] Step 3: Establish a multi-layer perceptron prediction model for predicting the algorithm's performance evaluation value, and use the multi-agent reinforcement learning algorithm evaluation sample in the actual offshore fishery supervision application scenario Train the prediction model, improve the loss function, and obtain better training results. Step 3 includes:

[0065] Step 3-1, create Figure 3 The multi-layer perceptron prediction model shown in the figure consists of three fully connected layers of different depths. The input layer contains 27-dimensional features, and then passes through two hidden layers to finally obtain the algorithm performance prediction value that imitates expert knowledge.

[0066] Step 3-2, The 2D evaluation samples are input into the prediction model, and feature transformation is performed through two fully connected layers with different depths. The output layer dimension is adjusted to a 2D tensor dimension, and a 1D algorithm evaluation prediction value is output to train the prediction model.

[0067] in, Indicates the number of evaluation samples used to train the model, 30 represents the corresponding evaluation samples The number of features in , including the target feature.

[0068] In step 3-3, the trained model can simulate an expert and predict the algorithm performance evaluation value of the sample to be evaluated. At the same time, the evaluation results are introduced into the loss function during the multi-layer perceptron training process to improve the loss function design and enhance the learning accuracy of the algorithm performance evaluation.

[0069] for samples , the detailed improvement design of the loss function is as follows:

[0070] (1)

[0071] in, is the predicted value of the algorithm performance evaluation for the corresponding evaluation sample, is the actual algorithm performance score of the corresponding evaluation sample, is the lower bound of the algorithm performance evaluation range of the corresponding evaluation sample, is the upper bound of the comprehensive evaluation range of the algorithm performance of the corresponding evaluation sample, represents the modified linear unit function.

[0072] Step 4: From samples Select appropriate data to form the background data set At this time, the background data set only contains 9 indicators and their upper and lower bounds. Based on the empirical conditional distribution modeling, the dependencies between features are formed in the background data set to form a conditional distribution that satisfies feature independence. Step 4 includes:

[0073] Step 4-1: Select 9 indicators and their upper and lower bounds from the evaluation sample, exclude other indicators, and samples Select appropriate data to form the background data set , Background dataset The total number of samples in .

[0074] Step 4-2, the mathematical definition of the empirical conditional distribution is as follows:

[0075] (2)

[0076] in, Indicates that it contains a subset of features A sample of Represents the complement of a feature subset samples; yes and The joint probability distribution of Indicates that the sample is in a known feature subset Under the condition of The conditional probability distribution of is the marginal probability density.

[0077] Step 4-3, in real data, when the feature dimension is high, and joint distribution is usually unknown, so the theoretical conditional distribution cannot be directly calculated, and the approximate empirical conditional distribution modeling dependencies between features:

[0078] (3)

[0079] Among them, the background dataset Satisfaction Among all samples of the condition, The distribution of is the conditional distribution, which is used to construct Approximation, That means approximation of To meet the background data The number of samples, is the complement of the feature subset No. samples; is the Dirac function, The pulses are added with equal weights, so that the original continuous probability distribution is transformed into a discrete probability distribution, and all sample points that meet the conditions are covered.

[0080] Step 5: For the evaluation prediction model obtained in the above step, calculate the gradient of the evaluation prediction model output with respect to the input. Introduce the empirical conditional distribution and model gradient into the Deep-SHAP method to obtain the improved Deep-SHAP interpretable method based on the empirical conditional distribution. Calculate the Shapley value of the algorithm performance evaluation prediction value of all indicators in the evaluation index system for the sample to be evaluated in the background data set. Step 5 includes:

[0081] In step 5-1, the algorithm performance evaluation of the prediction model outputs the gradient of the sample to be evaluated as follows:

[0082] (4)

[0083] in, Represents the model output prediction for a single sample The gradient is further decomposed into each specific feature , That is, the model output prediction for a single sample The gradient of the first feature in , and so on.

[0084] In step 5-2, the gradient is introduced into the Deep-SHAP method, and the Deep-SHAP calculation method for a single sample corrected by the model gradient is obtained as follows:

[0085] (5)

[0086] in, Indicates the Features in samples Contribution to the model's prediction output, Representation characteristics In the background dataset The mean value in Is the model prediction output relative to gradient.

[0087] In step 5-3, the approximate empirical conditional distribution is introduced into formula (5), and the improved Deep-SHAP calculation method for a single sample with the empirical conditional distribution is obtained as follows, which can calculate the Shapley value of the indicator to the model prediction output in a single sample to be evaluated:

[0088] (6)

[0089] in, Indicates that after the empirical conditional distribution is introduced into the Deep-SHAP method, Features in samples Contribution to the model's predictive output; Represents the approximate empirical conditional distribution The number of samples sampled in Indicates the Features in samples The gradient of is the same as that in formula (5), both of which are derived from formula (4); Representation characteristics In the conditions The approximate conditional expectation under .

[0090] In step 5-4, weighting all samples in the empirical conditional distribution yields the following improved Deep-SHAP calculation method based on the empirical conditional distribution. The Shapley value of the indicator for the predicted output in the background dataset can be calculated as follows:

[0091] (7)

[0092] in, Represents the entire background dataset Medium Features Contribution to the forecast output; Represents the approximate empirical conditional distribution The number of samples sampled in Indicates the Features in samples gradient.

[0093] Step 6: Based on the Shapley values ​​of the single sample and background dataset obtained above, visualize the feature importance ranking.

[0094] The present invention not only constructs a complete algorithm evaluation index system applicable to different application scenarios in a hierarchical manner; at the same time, it proposes a multi-agent reinforcement learning algorithm performance evaluation and prediction method based on a multi-layer perceptron, improves the loss function of learning and training, and obtains a high-precision prediction of the performance evaluation value of the multi-agent reinforcement learning algorithm to be evaluated; compared with traditional interpretation methods, it proposes a Deep-SHAP interpretation method based on improved empirical conditional distribution, which solves the problem of feature dependency in actual application scenarios, can achieve high-precision explainable prediction of multi-agent reinforcement learning algorithms, and obtain the feature importance ranking of multi-agent reinforcement learning algorithms.

[0095] The experimental results of the present invention are as follows Figure 4 The hardware environment is as follows: CPU i9-10 980XE, GPU RTX4070 Ti, 128 GB of memory; the software environment is Python 3.11.7, Pytorch 1.8.2. The original data of evaluation indicators for the simulation of a multi-agent reinforcement learning intelligent algorithm based on an offshore fishery supervision scenario was collected. Experts determined the corresponding algorithm evaluation values ​​based on the actual data performance of different algorithms, forming 100 evaluation samples. For the training of the multilayer perceptron neural network, 20% of the evaluation samples were retained as the test set. Figure 4 Demonstrates feature importance ranking using a multi-agent reinforcement learning algorithm in an offshore fishery supervision scenario. Figure 4The horizontal axis represents the Shapley value, which quantifies the contribution of each indicator to the model's single prediction output, ranging from approximately -500 to 20,000. The data to the left of each indicator represents the characteristic value of that indicator. A positive Shapley value indicates that the indicator has a positive contribution to the model's specific output, while a negative value indicates a negative contribution. The vertical axis lists the 27 different indicators involved in the analysis, which include 9 indicators and their upper and lower bounds, arranged in descending order by the average value of the feature importance. A red bar indicates that the feature made a positive contribution to the model output in that sample; conversely, a blue bar indicates that the feature made a negative contribution in that sample.

[0096] from Figure 4 It can be seen that the algorithm's memory usage is the top feature, and its overall impact on the model output is the greatest; the maximum algorithm memory usage, the maximum target ship's total voyage, the maximum job completion rate, and the job completion rate have a greater impact on the model's predicted output, while the other features have a smaller impact, verifying the present invention's ability to interpretably evaluate the performance of multi-agent reinforcement learning algorithms and visualize the importance ranking of features.

[0097] The features of the present invention include: the present invention comprehensively considers the principles of completeness, hierarchy and accuracy, and constructs a set of scientific multi-agent reinforcement learning algorithm performance evaluation index systems for multi-agent reinforcement learning algorithms in different practical application scenarios, providing a basis for the interpretability of the algorithm; the proposed algorithm performance evaluation prediction model based on multi-layer perceptron adopts real sample data training model, improves the loss function of learning and training, and obtains a high-precision prediction of the performance evaluation value of the multi-agent reinforcement learning algorithm to be evaluated; finally, in order to solve the problem of feature dependency in the actual application scenarios of the multi-agent reinforcement learning algorithm, the present invention proposes a Deep-SHAP interpretation method based on empirical conditional distribution improvement, which can achieve high-precision interpretable prediction of the multi-agent reinforcement learning algorithm and obtain the feature importance ranking of the multi-agent reinforcement learning algorithm.

[0098] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention. Any components not specified in this embodiment may be implemented using existing technologies.

Claims

1. A method for interpretable evaluation of the performance of multi-agent reinforcement learning algorithms based on Shapley additive interpretation, characterized by: The steps include: Step 1: Establish a simulation environment for a multi-agent reinforcement learning algorithm at sea and develop an evaluation index system suitable for the application scenarios of the multi-agent reinforcement learning algorithm. Specifically, in the offshore fishery supervision scenario, a hierarchical evaluation index system is constructed based on the algorithm's application in the actual scenario, including operational execution efficiency, operational and safety capabilities, repelling and interception capabilities, resource consumption, and algorithm performance. Step 2: Based on the maritime multi-agent reinforcement learning algorithm simulation environment and evaluation index system established in step 1, simulate the maritime multi-agent reinforcement learning algorithm simulation environment algorithm under different conditions in the maritime simulation environment, collect raw data of evaluation indicators, use the raw data of indicators to obtain algorithm evaluation values ​​based on expert knowledge, and form the collected raw data of indicators and algorithm evaluation values ​​into evaluation samples; Step 3: Build an evaluation prediction model based on a multi-layer perceptron to imitate expert knowledge and predict the algorithm performance evaluation value; train the evaluation prediction model based on the evaluation samples obtained in step 2, and improve the design of the loss function of the evaluation prediction model learning and training to obtain a better learning and training model to predict the algorithm performance evaluation value of the sample to be evaluated; Step 4: Based on the evaluation samples obtained in step 2, select data to form a background data set; model the dependencies between features based on the empirical conditional distribution, and form a conditional distribution that satisfies feature independence in the background data set; Step 5: Based on the empirical conditional distribution obtained in step 4, the empirical conditional distribution and model gradient are introduced into the Deep-SHAP method to obtain the improved Deep-SHAP interpretable method based on the empirical conditional distribution, thereby calculating the Shapley value of the predicted value of the algorithm performance evaluation of all indicators in the evaluation index system, including: In step 5-1, the algorithm performance evaluation of the prediction model outputs the gradient of the sample to be evaluated as follows: (4) in, Represents the model output prediction for a single sample The gradient is further decomposed into each specific feature , That is, the model output prediction for a single sample The gradient of the first feature in , and so on; In step 5-2, the gradient is introduced into the Deep-SHAP method, and the Deep-SHAP calculation method for a single sample corrected by the model gradient is obtained as follows: (5) in, Indicates the Features in samples Contribution to the model's prediction output, Representation characteristics In the background dataset The mean value in Is the model prediction output relative to gradient; In step 5-3, the approximate empirical conditional distribution is introduced into formula (5) to obtain the improved Deep-SHAP calculation method for a single sample of the empirical conditional distribution as follows, which calculates the Shapley value of the indicator on the model prediction output in a single sample to be evaluated: (6) in, Indicates that after the empirical conditional distribution is introduced into the Deep-SHAP method, Features in samples Contribution to the model's predictive output; Indicates the Features in samples The gradient of is the same as that in formula (5), both of which are derived from formula (4); Representation characteristics In the conditions approximate conditional expectation under ; In step 5-4, weighting all samples in the empirical conditional distribution yields the following improved Deep-SHAP calculation method based on the empirical conditional distribution. The Shapley value of the indicator pair prediction output in the background dataset is calculated as follows: (7) in, Represents the entire background dataset Medium Features Contribution to the prediction output, Represents the approximate empirical conditional distribution The number of samples sampled in Step 6: Visualize the feature importance ranking based on the Shapley value obtained in step 5.

2. The method for interpretable evaluation of the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation according to claim 1, characterized in that: Step 2 includes: Based on the constructed hierarchical evaluation index system and the marine multi-agent reinforcement learning algorithm simulation environment, the multi-agent reinforcement learning algorithms under different application scenarios are simulated in the marine environment, and the original data of the evaluation indicators are collected. The upper and lower bounds of the indicator data are set according to the historical data of different algorithms. The weights are calculated according to the standardization of the indicator data, and the scores of each indicator are weighted to obtain the corresponding algorithm evaluation value. The collected original indicator data and algorithm evaluation values ​​of the evaluation samples constitute the evaluation samples.

3. The method for interpretable evaluation of the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation according to claim 2, characterized in that: Step 3 includes: Step 3-1: Build a multi-layer perceptron prediction model. The prediction model consists of three fully connected layers with different depths. The input layer contains 27-dimensional features, and then passes through two hidden layers to finally obtain the algorithm performance prediction value that simulates expert knowledge. Step 3-2, The 2D evaluation samples are input into the prediction model, and feature transformation is performed through two fully connected layers with different depths. The output layer dimension is adjusted to a 2D tensor dimension, and the algorithm evaluation prediction value is output to train the prediction model. in, Indicates the number of evaluation samples used to train the model, 30 represents the corresponding evaluation samples The number of features in ; Step 3-3: Use the trained model to simulate experts and predict the algorithm performance evaluation value of the samples to be evaluated. At the same time, introduce the evaluation results into the loss function of the multi-layer perceptron training process to improve the design of the loss function. for samples , the improved design of the loss function is as follows: (1) in, is the predicted value of the algorithm performance evaluation for the corresponding evaluation sample, is the actual algorithm performance score of the corresponding evaluation sample, is the lower bound of the algorithm performance evaluation range of the corresponding evaluation sample, is the upper bound of the comprehensive evaluation range of the algorithm performance of the corresponding evaluation sample, represents the rectified linear unit function.

4. The method for interpretable evaluation of the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation according to claim 3, characterized in that: Step 4 includes: Step 4-1: Select 9 indicators and their upper and lower bounds from the evaluation sample. The 9 indicators include operation completion time, operation completion rate, patrol boat group range, patrol boat group online rate, algorithm decision time, algorithm memory usage, target ship total range, illegal zone approach threshold and equipment minimum safety distance. Eliminate other indicators and select samples Select data from the background dataset , Background dataset The total number of samples in ; Step 4-2, the mathematical definition of the empirical conditional distribution is as follows: (2) in, Indicates that it contains a subset of features A sample of Represents the complement of a feature subset samples; yes and The joint probability distribution of Indicates that the sample is in a known feature subset Under the condition of The conditional probability distribution of is the marginal probability density; Step 4-3, approximate the empirical conditional distribution to model the dependencies between features: (3) Among them, the background dataset Satisfaction Among all samples of the condition, The distribution of is the conditional distribution, which is used to construct Approximation, That means approximation of To meet the background data The number of samples, is the complement of the feature subset No. samples; is the Dirac function, The pulses are added with equal weights, so that the original continuous probability distribution is transformed into a discrete probability distribution, and all sample points that meet the conditions are covered.

5. The method for interpretable evaluation of the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation according to claim 1, characterized in that: In step 5, the feature importance ranking is visualized based on the obtained Shapley values ​​of the single sample and the background dataset.

6. A multi-agent reinforcement learning algorithm performance interpretable evaluation system based on Shapley additive interpretation, characterized by: For implementing the method described in any one of claims 1 to 5, the system comprises: The first module is used to establish a marine algorithm simulation environment and evaluation index system that is suitable for the application scenarios of multi-agent reinforcement learning algorithms; The second module, based on the algorithm simulation environment and evaluation index system established in the first module, simulates the multi-agent reinforcement learning algorithm in different actual scenarios in the marine algorithm simulation environment, collects the original data of the evaluation indicators, and uses the original data of the indicators to obtain the algorithm evaluation value based on expert knowledge. The collected original data of the indicators and the algorithm evaluation value constitute the evaluation sample; The third module builds an evaluation prediction model based on a multi-layer perceptron, imitating expert knowledge to predict the algorithm performance evaluation value; the evaluation prediction model is trained based on the evaluation samples obtained in the second module, and the loss function design of the evaluation prediction model learning and training is improved to obtain a better learning and training model to predict the algorithm performance evaluation value of the samples to be evaluated; The fourth module selects data to form a background data set based on the evaluation samples obtained in the second module, models the dependencies between features based on the empirical conditional distribution, and forms a conditional distribution that satisfies feature independence in the background data set; The fifth module calculates the Shapley value of the algorithm performance evaluation prediction value of all indicators in the evaluation index system for the samples to be evaluated based on the background data set obtained in the fourth module and the improved Deep-SHAP method; The sixth module visualizes the feature importance ranking based on the Shapley value obtained in the fifth module.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the interpretable evaluation method for the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the interpretable evaluation method for the effectiveness of a multi-agent reinforcement learning algorithm based on Shapley additive interpretation as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning system and method, electronic equipment and storage medium

    CN117933350A

  • Rapid Shapley value estimation method based on stratified sampling

    CN119047589A