Multidimensional index-based reinforcement learning algorithm performance evaluation method, device and equipment
Through comprehensive analysis of multi-dimensional indicators, the problem of single indicators in the performance evaluation of reinforcement learning algorithms was solved, and a comprehensive evaluation and optimization of the algorithm's efficiency, stability and security were achieved, improving the algorithm's performance in complex environments.
Patent Information
- Application Number
- CN202510677195.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-10-17
AI Technical Summary
Existing methods for evaluating the performance of reinforcement learning algorithms rely on a single metric, which makes it difficult to fully reflect the overall performance of the algorithm. In particular, in complex and dynamic environments, they cannot accurately measure the impact of factors such as the balance between exploration and exploitation, reward design, and environmental characteristics.
A comprehensive analysis method using multidimensional indicators is adopted. By collecting and preprocessing multidimensional indicator data during reinforcement learning training, efficiency, stability, and security scores are calculated. Pearson correlation coefficient is used to analyze the correlation between indicators and generate optimization suggestions to adjust the balance between exploration and exploitation and the reward function.
It achieves more accurate and comprehensive performance evaluation of reinforcement learning algorithms, improves the efficiency, stability and security of the algorithms in complex environments, and enhances the robustness and generalization ability of the algorithms.
Smart Images

Figure CN120803859A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of reinforcement learning, in particular to a performance evaluation method and device for reinforcement learning algorithm based on multi-dimensional indicators and equipment. BACKGROUND
[0002] Reinforcement Learning (RL) is a machine learning method that learns strategies through trial and error with the environment. The core goal is to maximize cumulative rewards to guide model training to obtain optimal decisions through continuous experimentation. In recent years, reinforcement learning has been widely applied in autonomous driving, robot control, game AI, financial decision-making and other fields, showing great potential.
[0003] However, the existing performance evaluation methods for reinforcement learning algorithms still have many challenges. The performance evaluation of reinforcement learning algorithms usually relies on a single indicator or a few performance dimensions, such as cumulative rewards, policy entropy, or sampling efficiency. Although these indicators can reflect the performance of the algorithm in some aspects, they cannot fully characterize the overall performance of the algorithm. In practical applications, reinforcement learning algorithms need to learn in complex and dynamic environments, and their performance is affected by various factors such as the trade-off between exploration and exploitation, reward design, and environmental characteristics. In addition, a large number of multi-dimensional indicators are generated during training, such as value functions (Q-values), policy entropy, and gradient changes. These indicators have complex interdependent relationships. Therefore, relying on a single indicator or a few indicators for evaluation cannot accurately reflect the overall performance of the algorithm, and needs to be addressed. SUMMARY
[0004] The present application provides a performance evaluation method and device for reinforcement learning algorithm based on multi-dimensional indicators to solve the problem of single indicator in existing reinforcement learning evaluation methods and the difficulty of accurately reflecting the overall performance of the algorithm. Through comprehensive analysis of multi-dimensional indicators, more accurate and comprehensive evaluation and optimization are achieved.
[0005] The first aspect embodiment of the present application provides a performance evaluation method for reinforcement learning algorithm based on multi-dimensional indicators, comprising the following steps:
[0006] Collecting multi-dimensional indicator data during reinforcement learning training and preprocessing the multi-dimensional indicator data to obtain preprocessed multi-dimensional indicator data;
[0007] Calculating efficiency score, stability score and safety score based on the preprocessed multi-dimensional indicator data, and analyzing the correlation between multi-dimensional indicators based on the efficiency score, the stability score and the safety score according to a preset Pearson correlation coefficient calculation formula to obtain a correlation evaluation result;
[0008] generate an optimization suggestion strategy based on the correlation evaluation result, and perform an optimization action according to the optimization suggestion strategy.
[0009] According to an embodiment of the present application, the generating an optimization suggestion strategy based on the correlation evaluation result, and performing an optimization action according to the optimization suggestion strategy comprises:
[0010] obtaining a current policy entropy value, a current gradient update frequency, a current reward signal and a current cumulative return;
[0011] based on the correlation evaluation result, adjusting a balance strategy between exploration and utilization according to the current policy entropy value and the current gradient update frequency, and / or reconstructing a current reward function according to the correlation between the current reward signal and the current cumulative return.
[0012] According to an embodiment of the present application, after calculating the efficiency score, the stability score and the safety score based on the preprocessed multi-dimensional index data, it further comprises:
[0013] obtaining a comprehensive score according to the efficiency score, the stability score and the safety score;
[0014] wherein the comprehensive score is:
[0015] S total = w1S efficiency + w2S stability + w3S safety
[0016] wherein S total is the comprehensive score, w1 is the weight of the comprehensive score, S efficiency is the comprehensive score, w2 is the weight of the stability score, S stability is the stability score, and w3 is the weight of the safety score, S safety is the safety score.
[0017] According to an embodiment of the present application, the performance evaluation method of the multi-dimensional index based reinforcement learning algorithm further comprises:
[0018] based on a preset time series curve graph, monitoring the change trend of each index in the multi-dimensional index data with time steps, and visually displaying the comprehensive score.
[0019] According to an embodiment of the present application, the preset Pearson correlation coefficient calculation formula is:
[0020]
[0021] wherein r xyis a correlation coefficient of the x variable and the y variable, n is a number of observation values, i is an i-th observation value, x i is an i-th x variable, is a mean value of the variable x, y i is an i-th y variable, is a mean value of the variable y.
[0022] According to an embodiment of the present application, the multi-dimensional index data includes at least one of a Critic-related index, a Policy-related index, environment feedback data, and a reward signal; the Critic-related index includes at least one of an average Q value, a first standard deviation, and a first gradient norm, and the Policy-related index includes at least one of a policy mean value, a second standard deviation, an entropy value, and a second gradient norm.
[0023] According to the multi-dimensional index-based reinforcement learning algorithm performance evaluation method provided by the embodiments of the present application, the multi-dimensional index data in the reinforcement learning training process is preprocessed, the efficiency score, the stability score, and the safety score are calculated based on the preprocessed multi-dimensional index data, the correlation between the multi-dimensional indexes is analyzed based on the preset Pearson correlation coefficient calculation formula according to the efficiency score, the stability score, and the safety score, the correlation evaluation result is obtained, and the optimization suggestion strategy is generated to perform the optimization action according to the optimization suggestion strategy. Thus, the problem of single index in the existing reinforcement learning evaluation method and the difficulty in accurately reflecting the overall performance of the algorithm are solved, and more accurate and comprehensive evaluation and optimization are achieved through comprehensive analysis of multi-dimensional indexes.
[0024] The second aspect embodiment of the present application provides a multi-dimensional index-based reinforcement learning algorithm performance evaluation device, which comprises:
[0025] The acquisition module is configured to acquire multi-dimensional index data in a reinforcement learning training process, and pre-process the multi-dimensional index data to obtain pre-processed multi-dimensional index data.
[0026] The calculation module is configured to calculate an efficiency score, a stability score, and a safety score based on the pre-processed multi-dimensional index data, and analyze the correlation between the multi-dimensional indexes based on a preset Pearson correlation coefficient calculation formula according to the efficiency score, the stability score, and the safety score, to obtain a correlation evaluation result.
[0027] The generation module is configured to generate an optimization suggestion strategy based on the correlation evaluation result, and perform an optimization action according to the optimization suggestion strategy.
[0028] According to an embodiment of the present application, the generation module is configured to:
[0029] obtain a current policy entropy value, a current gradient update frequency, a current reward signal and a current cumulative return;
[0030] based on the correlation evaluation result, adjust a balance strategy between exploration and utilization according to the current policy entropy value and the current gradient update frequency, and / or reconstruct a current reward function according to the correlation between the current reward signal and the current cumulative return.
[0031] According to an embodiment of the present application, after calculating the efficiency score, the stability score and the safety score based on the preprocessed multi-dimensional index data, the computing module is further configured to:
[0032] obtain a comprehensive score based on the efficiency score, the stability score and the safety score;
[0033] wherein the comprehensive score is:
[0034] S total = w1S efficiency + w2S stability + w3S safety
[0035] wherein S total is the comprehensive score, w1 is the weight of the comprehensive score, S efficiency is the comprehensive score, w2 is the weight of the stability score, S stability is the stability score, and w3 is the weight of the safety score, S safety is the safety score.
[0036] According to an embodiment of the present application, the multi-dimensional index-based reinforcement learning algorithm performance evaluation device is further configured to:
[0037] based on a preset time series curve graph, monitor the change trend of each index in the multi-dimensional index data with time steps, and visually display the comprehensive score.
[0038] According to an embodiment of the present application, the preset Pearson correlation coefficient calculation formula is:
[0039]
[0040] wherein r xy is the correlation coefficient of the x variable and the y variable, n is the number of observation values, i is the i-th observation value, x i is the i-th x variable, is the mean of the variable x, y i is the i-th y variable, is the mean of the variable y.
[0041] According to an embodiment of the present application, the multi-dimensional index data comprises at least one of Critic-related indexes, Policy-related indexes, environment feedback data and reward signals; the Critic-related indexes comprise at least one of an average Q value, a first standard deviation and a first gradient norm, and the Policy-related indexes comprise at least one of a policy mean value, a second standard deviation, an entropy value and a second gradient norm.
[0042] According to the multi-dimensional index-based reinforcement learning algorithm performance evaluation device provided by the embodiment of the present application, the multi-dimensional index data in the reinforcement learning training process is preprocessed, the efficiency score, the stability score and the safety score are calculated based on the preprocessed multi-dimensional index data, the correlation between the multi-dimensional indexes is analyzed based on the preset Pearson correlation coefficient calculation formula according to the efficiency score, the stability score and the safety score, the correlation evaluation result is obtained, and the optimization suggestion strategy is generated to perform the optimization action according to the optimization suggestion strategy. Thus, the problem of single index and difficulty in accurately reflecting the overall performance of the algorithm in the existing reinforcement learning evaluation method is solved, and more accurate and comprehensive evaluation and optimization are realized through comprehensive analysis of multi-dimensional indexes.
[0043] The third aspect embodiment of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the multi-dimensional index-based reinforcement learning algorithm performance evaluation method as described in the above embodiments.
[0044] The fourth aspect embodiment of the present application provides a computer readable storage medium, which stores computer instructions for making the computer execute the multi-dimensional index-based reinforcement learning algorithm performance evaluation method as described in the above embodiments.
[0045] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0046] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings, wherein:
[0047] Figure 1 A flowchart of a multi-dimensional index-based reinforcement learning algorithm performance evaluation method according to an embodiment of the present application;
[0048] Figure 2 A block diagram of a multi-dimensional index-based reinforcement learning algorithm performance evaluation system according to an embodiment of the present application;
[0049] Figure 3 A block schematic diagram of a performance evaluation device for a multi-dimensional index-based reinforcement learning algorithm according to an embodiment of the present application;
[0050] Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0051] Embodiments of the present application are described in detail below with reference to examples shown in the attached drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.
[0052] A multi-dimensional index-based reinforcement learning algorithm performance evaluation method, device and equipment of an embodiment of the present application are described below with reference to the drawings. In view of the problem of single index and difficulty in accurately reflecting the overall performance of the algorithm in the existing reinforcement learning evaluation method mentioned in the background art, the present application provides a multi-dimensional index-based reinforcement learning algorithm performance evaluation method. In the method, the multi-dimensional index data in the reinforcement learning training process is preprocessed, the efficiency score, stability score and safety score are calculated based on the preprocessed multi-dimensional index data, and the correlation between the multi-dimensional indexes is analyzed based on the preset Pearson correlation coefficient calculation formula according to the efficiency score, stability score and safety score to obtain a correlation evaluation result, and an optimization suggestion strategy is generated to perform an optimization action according to the optimization suggestion strategy. Thus, the problem of single index and difficulty in accurately reflecting the overall performance of the algorithm in the existing reinforcement learning evaluation method is solved, and more accurate and comprehensive evaluation and optimization are achieved through comprehensive analysis of multi-dimensional indexes.
[0053] Specifically, Figure 1 A flowchart of a multi-dimensional index-based reinforcement learning algorithm performance evaluation method provided by an embodiment of the present application.
[0054] As Figure 1 shown, the multi-dimensional index-based reinforcement learning algorithm performance evaluation method includes the following steps:
[0055] In step S101, multi-dimensional index data in the reinforcement learning training process is collected, and the multi-dimensional index data is preprocessed to obtain preprocessed multi-dimensional index data.
[0056] In some embodiments, the multi-dimensional indicator data includes at least one of Critic-related indicators, Policy-related indicators, environment feedback data, and reward signals; the Critic-related indicators include at least one of an average Q value, a first standard deviation, and a first gradient norm; and the Policy-related indicators include at least one of a policy mean, a second standard deviation, an entropy value, and a second gradient norm.
[0057] Specifically, the embodiments of the present application can collect multi-dimensional indicator data generated in real time during the reinforcement learning training process through a sampler, which can include Critic-related indicators, Policy-related indicators, environment feedback data, and reward signals.
[0058] The Critic-related indicators include an average Q value, a first standard deviation, and a first gradient norm. The average Q value (critic_avg_q1 / q2) is used to measure the estimation level of the value function. The first standard deviation (std1 / std2) is used to reflect the stability of the value function estimation. The first gradient norm (q_grad_norm) represents the magnitude of the value function update, and the calculation formula is as follows:
[0059]
[0060] wherein L Q is the loss function of the value function, and θ Q is the parameter of the value function.
[0061] Further, the Policy-related indicators can include a policy mean, a second standard deviation, an entropy value, and a second gradient norm. The policy mean and the standard deviation are used to represent the action distribution output by the policy network, the entropy value (entropy) is used to evaluate the randomness of the exploration behavior, and the second gradient norm (policy_grad_norm) is used to measure the magnitude of the policy network update. The calculation formula of the entropy value is as follows:
[0062]
[0063] wherein H(π) is the entropy value, used to measure the randomness of the policy, a is the action taken by the current policy, and s is the current environment state. The higher the entropy value, the more active the exploration behavior.
[0064] Further, the environment feedback data can include TAR cumulative return, task end reason, and violation rate. The TAR cumulative return is used to measure the overall task completion of the agent. The task end reason includes the distribution of events such as completing the target, timeout, and interruption. The violation rate mainly includes some specific indicators, such as the out-of-bound rate under certain constraint conditions, etc.
[0065] Further, the reward signal can include dimensional tracking errors and control costs. The dimensional tracking errors are used to measure the error between the agent and the target value, such as longitudinal error and lateral error. The control costs are used to measure the punishment of the agent performing actions beyond the boundary, such as the punishment of the direction turning angle.
[0066] Further, the multi-dimensional index data is standardized and pre-processed, such as normalized reward signal, to ensure data comparability. In addition, the mean, standard deviation and the like are calculated for part of the indicators, so as to analyze the stability of the algorithm.
[0067] For example, the normalization processing of the embodiment of the application can be scaling of the reward signal, gradient norm and the like, and the formula is as follows:
[0068]
[0069] Wherein, x' is the normalized value, x is the original value to be normalized.
[0070] Further, the statistical characteristics are extracted, and the policy mean and standard deviation are calculated for evaluating the stability, for example:
[0071]
[0072] Wherein, Mean is the mean, n is the sampling number, x i is the actual value of the sample, and Std is the standard deviation.
[0073] In step S102, the efficiency score, the stability score and the safety score are calculated based on the pre-processed multi-dimensional index data, and the correlation between the multi-dimensional indexes is analyzed based on the preset Pearson correlation coefficient calculation formula according to the efficiency score, the stability score and the safety score, to obtain the correlation evaluation result.
[0074] Specifically, the efficiency score of the embodiment of the application is calculated based on the cumulative return and the algorithm training time, and the formula is as follows:
[0075]
[0076] Wherein, S efficiency is the efficiency score, TAR is the cumulative return, and Training Time is the algorithm training time.
[0077] Further, the stability score is calculated based on the change range of the Q value and the policy gradient, and the formula is as follows:
[0078]
[0079] Wherein, S stability is the stability score, a standard deviation of the gradient of the value function, a maximum value of the gradient of the value function.
[0080] Further, the safety score is obtained based on a violation rate, a collision rate and the like of the strategy in the environment, and the formula is as follows:
[0081] S safety =1-(Violation Rate+Collision Rate)
[0082] wherein S safety is the safety score, the Violation Rate is the violation rate of the strategy in the environment, and the Collision Rate is the collision rate of the strategy in the environment.
[0083] Further, in some embodiments, after the efficiency score, the stability score and the safety score are calculated based on the preprocessed multi-dimensional index data, the method further comprises: calculating a comprehensive score according to the efficiency score, the stability score and the safety score and the weights of different dimensions; wherein the comprehensive score is:
[0084] S total =w1S efficiency +w2S stability +w3S safety
[0085] wherein S total is the comprehensive score, w1 is the weight of the comprehensive score, S efficiency is the comprehensive score, w2 is the weight of the stability score, S stability is the stability score, and w3 is the weight of the safety score, S safety is the safety score.
[0086] Further, based on a preset Pearson correlation coefficient calculation formula, the correlation between the indexes is analyzed according to the efficiency score, the stability score and the safety score. The user can determine the key influencing factors by using the Pearson correlation coefficient to analyze the correlation between the indexes. The preset Pearson correlation coefficient calculation formula is:
[0087]
[0088] wherein r xy is the correlation coefficient of the x variable and the y variable, n is the number of observation values, i is the i-th observation value, x i is the i-th x variable, is the mean of the variable x, y i is the i-th y variable, is the mean of the variable y.
[0089] In step S103, an optimization suggestion strategy is generated based on the correlation evaluation result, so as to perform an optimization action according to the optimization suggestion strategy.
[0090] Further, in some embodiments, the optimization suggestion strategy is generated based on the correlation evaluation result, so as to perform an optimization action according to the optimization suggestion strategy, including: obtaining a current policy entropy value, a current gradient update frequency, a current reward signal and a current cumulative return; based on the correlation evaluation result, adjusting a balance strategy between exploration and exploitation according to the current policy entropy value and the current gradient update frequency, and / or reconstructing a current reward function according to the correlation between the current reward signal and the current cumulative return.
[0091] Specifically, in reinforcement learning, an agent needs to find a balance between exploration and exploitation. Exploration refers to the agent trying new actions to obtain more information about the environment, while exploitation refers to the agent selecting the optimal action based on existing knowledge to maximize cumulative rewards. Policy entropy is an indicator of policy randomness. The higher the entropy value, the more random the policy, and the more active the exploration behavior; the lower the entropy value, the more deterministic the policy, and the more obvious the exploitation behavior. Gradient update frequency refers to the frequency of updating the policy network during training.
[0092] Further, based on the policy entropy and the gradient update frequency, the following optimization suggestions can be proposed: if the policy entropy is low, it means that the policy is too deterministic, and there is a problem of over-exploitation. At this time, exploration can be increased to avoid local optimal solution. If the gradient update frequency is too high, it will lead to unstable training process. By using larger batch data for updating, the update frequency can be reduced, and the training stability can be improved.
[0093] As can be appreciated by those skilled in the art, the reward signal is an important factor in reinforcement learning to guide the behavior of the agent. The cumulative return is the total reward obtained by the agent in a period of time, reflecting the overall performance of the agent. The reward signal is the immediate feedback of the environment to the agent's behavior. For example, in the autonomous driving scenario, the reward signal can include vehicle speed, position error, safe distance from other vehicles, etc. The cumulative return is the total reward obtained by the agent in multiple time steps, reflecting the performance of the agent in long-term tasks.
[0094] Based on the correlation between the reward signal and the cumulative return, the following optimization suggestions can be proposed: if certain reward signals are highly correlated with the cumulative return, it indicates that these signals have an important impact on the long-term performance of the agent, and the weight of these signals in the reward function can be increased to make the agent pay more attention to these signals. If the value range of certain reward signals is too large or too small, it will affect the stability of the training process. The range of the reward signal can be normalized or adjusted to make its impact on the training process more reasonable.
[0095] Further, in some embodiments, the multi-dimensional index-based reinforcement learning algorithm performance evaluation method further comprises: based on a preset time series graph, monitoring the change trend of each index in the multi-dimensional index data with time steps, and visualizing the comprehensive score.
[0096] Specifically, the embodiments of the present application can call the Tensorboard service to monitor the changes of various indicators with time steps during the training process, and realize real-time visualization of multi-dimensional data through custom front-end UI, including the following two aspects:
[0097] (1) Index trend monitoring: using a time series graph to show the change trend of various indicators (such as Q value, policy entropy, violation rate, etc.) with time steps during the training process.
[0098] (2) Algorithm performance comparison: visualizing the comprehensive score, which can show the comprehensive performance of the algorithm through a radar chart or a column chart, and then form a comparison chart of the algorithm performance in different scenarios.
[0099] Therefore, the multi-dimensional index-based reinforcement learning evaluation and optimization method proposed in the present application covers three major dimensions of algorithm efficiency and stability, environment reward, loss function and time consumption. Through the collection and analysis of multi-dimensional indicators, the performance of the reinforcement learning algorithm can be comprehensively reflected, including efficiency, stability, safety, environmental adaptability, etc., which makes up for the limitations of single indicator evaluation in the prior art. In addition, in diversified scenarios, the robustness and generalization ability of the algorithm are one of the main factors determining its actual application value. The present application refines the analysis of environmental evaluation indicators, such as violation rate and task completion efficiency in different scenarios, which helps to improve the robustness and generalization ability of the algorithm. The various indicators can also be monitored through a visual interface, reducing the manual intervention and time consumption of algorithm development and model debugging.
[0100] In order to make the person skilled in the art more clearly and intuitively understand the multi-dimensional index-based reinforcement learning algorithm performance evaluation and model optimization method proposed in the present application, the following will be described in detail with specific embodiments.
[0101] This example takes training multi-lane and intersection car automatic driving strategy as an example to show the actual application scenario and its key steps. The implementation process is as follows:
[0102] First, system initialization is performed, and environment configuration is completed. The simulation platform selects mature automatic driving simulation platforms such as CARLA or SUMO to simulate real multi-lane roads and intersection scenarios. Design a reward function suitable for task requirements, including but not limited to vehicle speed, position error, safe distance from other vehicles, and compliance with traffic rules, etc., and complete the setting of RL algorithm parameters.
[0103] Second, use the sampler in the reinforcement learning training tool deployed in the simulation loop to collect the following indicators in real time:
[0104] (1) Critic-related indicators: average Q value, standard deviation, gradient norm;
[0105] (2) Policy-related indicators: policy mean, entropy, gradient norm
[0106] (3) Environment feedback-related indicators: TAR cumulative return, task completion reason (record the specific circumstances of each task completion, such as whether it successfully reached the destination, whether there was a collision accident, etc.), violation rate (statistical number of traffic law violations, especially red light running, not maintaining a safe distance, etc.)
[0107] (4) Reward signal-related indicators: each dimension tracking error (calculate the deviation between the actual vehicle trajectory and the ideal path to help adjust the driving strategy), control cost (give appropriate punishment for operations that exceed the safe range, such as sudden braking or sudden lane changes to enhance safety.)
[0108] Further, the data of each time step is transmitted to the data processing module to perform the following operations:
[0109] (1) Normalization: unify data from different dimensions to the same numerical range for subsequent analysis and comparison. All reward signals are scaled to [-1, 1].
[0110] (2) Statistical property extraction: calculate the mean and standard deviation of Critic-related indicators and Policy-related indicators to provide a basis for subsequent stability and efficiency scoring.
[0111] Finally, after preprocessing and standardizing some indicators, perform index analysis and optimization, and observe through the visualization module:
[0112] Comprehensive score generation: Weights are assigned based on the importance of different dimensions to calculate the overall score. For example, when driving on city roads, safety may be given more weight; on highways, efficiency may be given a higher weight. In this example, the scenario involves multiple lanes and an intersection, so safety is given the highest weight.
[0113] Correlation Analysis: Use methods such as the Pearson correlation coefficient to analyze the relationships between scores and identify key factors influencing the final result. We've found that higher entropy values often lead to better exploration results. In this example, a more reasonable entropy value is required for the multi-lane scenario to help find a more optimal driving route.
[0114] Optimization suggestion generation: Based on the analysis results, specific improvement measures are proposed. For multi-lane scenarios, the exploration / exploitation balance is adjusted. For intersection scenarios, a higher weight is given to the penalty factor for violations in the reward design to ensure that the strategy is relatively conservative and further improve its performance.
[0115] Visual indicator monitoring: By manually setting parameters such as the number of training rounds, strategy saving interval size, and data processing step size, the Tensorboard tool is called to draw a time series curve chart to show the changing trend of each indicator over time during the training process, intuitively showing the progress of strategy learning and whether the strategy has converged after training.
[0116] Therefore, the embodiment of the present application can comprehensively reflect the performance of the reinforcement learning algorithm by covering a multi-dimensional evaluation system of critic-related indicators, policy-related indicators, environmental feedback data, and reward signals. At the same time, multiple indicators such as gradient norm, policy entropy, violation rate, etc. are introduced into the evaluation, taking into account the efficiency, stability and security of the algorithm. Indicator normalization and statistical characteristic extraction methods are designed, which are both efficient and accurate when processing data of different dimensions. An innovative method of calculating the total score based on the weighted efficiency score, stability score and security score is proposed, and the weights are flexibly adjusted to meet the needs of different tasks. In addition, feedback optimization suggestions can be generated in combination with the correlation analysis between indicators to guide the parameter adjustment of the reinforcement learning algorithm, and a real-time visualization interface for indicator data is designed. The intuitive comparison of indicator trends and algorithm performance is achieved through time series curves, which enhances the monitoring of the training process and the guidance of algorithm improvement.
[0117] The following describes a multi-dimensional RL algorithm performance evaluation method. This system includes a multi-dimensional RL algorithm performance evaluation system. This system features common metrics for numerous RL algorithms, supports a variety of RL algorithms, and is adaptable to different algorithmic frameworks. Furthermore, to enhance the system's scalability in evaluation scenarios, it allows users to extend it with custom metrics and dynamically adjust the weights of these metrics to accommodate different task requirements (e.g., prioritizing efficiency or safety).
[0118] Specifically, as shown in the figure, the multi-dimensional index-based reinforcement learning algorithm performance evaluation system architecture is composed of a data acquisition module, a data processing module, an index analysis and optimization module, and a visualization module. Figure 2
[0119] Among them, the data acquisition module is the input source of the system, which collects the algorithm indicators generated in real time during the training process through the sampler sampler. The data processing module mainly processes different dimension indicators, and standardizes and preprocesses the input indicators. The index analysis and optimization module is used for algorithm evaluation and optimization based on multi-dimensional indexes, including comprehensive score generation, correlation analysis, and feedback optimization suggestions. The visualization module is used to call the Tensorboard service to monitor the changes of various indicators with time steps during the training process, and to realize real-time visualization of multi-dimensional data through custom front-end UI.
[0120] Therefore, the multi-dimensional index system for reinforcement learning algorithm performance evaluation and model training effect evaluation provided by the present application covers algorithm utility evaluation indicators, environment reward evaluation indicators, loss function and time consumption evaluation indicators, and other multi-level dimensions. The system guides decision optimization and comprehensively evaluates the efficiency, stability and applicability of the algorithm through sampling and environment interaction data, normalization processing and multi-index fusion analysis, providing theoretical support and optimization path for reinforcement learning algorithm development and deployment and decision model training optimization.
[0121] The multi-dimensional index-based reinforcement learning algorithm performance evaluation method provided by the embodiment of the present application preprocesses the multi-dimensional index data in the reinforcement learning training process, calculates the efficiency score, stability score and safety score based on the preprocessed multi-dimensional index data, analyzes the correlation between the multi-dimensional indexes based on the preset Pearson correlation coefficient calculation formula according to the efficiency score, stability score and safety score, obtains the correlation evaluation result, and generates an optimization suggestion strategy, so as to execute the optimization action according to the optimization suggestion strategy. Therefore, the problem of single index in the existing reinforcement learning evaluation method and the difficulty in accurately reflecting the overall performance of the algorithm are solved, and more accurate and comprehensive evaluation and optimization are realized through comprehensive analysis of multi-dimensional indexes.
[0122] Next, the multi-dimensional index-based reinforcement learning algorithm performance evaluation device provided by the embodiment of the present application is described with reference to the accompanying drawings.
[0123] Figure 3 is a block schematic diagram of the multi-dimensional index-based reinforcement learning algorithm performance evaluation device of the embodiment of the present application.
[0124] As shown in the figure, the multi-dimensional index-based reinforcement learning algorithm performance evaluation device is composed of a data acquisition module, a data processing module, an index analysis and optimization module, and a visualization module. Figure 3 As shown, the multi-dimensional index-based reinforcement learning algorithm performance evaluation device 10 comprises a collection module 100, a calculation module 200, and a generation module 300.
[0125] The collection module 100 is configured to collect multi-dimensional index data in a reinforcement learning training process, and pre-process the multi-dimensional index data to obtain pre-processed multi-dimensional index data. The calculation module 200 is configured to calculate efficiency scores, stability scores, and safety scores based on the pre-processed multi-dimensional index data, and analyze the correlation between the multi-dimensional indexes based on a preset Pearson correlation coefficient calculation formula according to the efficiency scores, the stability scores, and the safety scores to obtain a correlation evaluation result. The generation module 300 is configured to generate an optimization suggestion strategy based on the correlation evaluation result, so as to perform an optimization action according to the optimization suggestion strategy.
[0126] Further, in some embodiments, the generation module 300 is configured to: obtain a current policy entropy value, a current gradient update frequency, a current reward signal, and a current cumulative return; adjust a balance strategy between exploration and utilization according to the current policy entropy value and the current gradient update frequency based on the correlation evaluation result, and / or reconstruct a current reward function according to the correlation between the current reward signal and the current cumulative return.
[0127] Further, in some embodiments, after calculating the efficiency scores, the stability scores, and the safety scores based on the pre-processed multi-dimensional index data, the calculation module 200 is further configured to: obtain a comprehensive score according to the efficiency scores, the stability scores, and the safety scores; wherein the comprehensive score is:
[0128] S total = w1S efficiency + w2S stability + w3S safety
[0129] wherein S total is the comprehensive score, w1 is the weight of the comprehensive score, S efficiency is the comprehensive score, w2 is the weight of the stability score, S stability is the stability score, and w3 is the weight of the safety score, S safety is the safety score.
[0130] Further, in some embodiments, the multi-dimensional index-based reinforcement learning algorithm performance evaluation device 10 is further configured to: based on a preset time series curve graph, monitor the change trend of each index in the multi-dimensional index data with time steps, and visually display the comprehensive score.
[0131] Further, in some embodiments, the preset Pearson correlation coefficient calculation formula is:
[0132]
[0133] wherein r xy is a correlation coefficient of the x variable and the y variable, n is the number of observation values, i is the i-th observation value, x i is the i-th x variable, is the mean of the variable x, y i is the i-th y variable, is the mean of the variable y.
[0134] Further, in some embodiments, the multi-dimensional index data includes at least one of Critic-related indexes, Policy-related indexes, environment feedback data and reward signals; the Critic-related indexes include at least one of an average Q value, a first standard deviation and a first gradient norm, and the Policy-related indexes include at least one of a policy mean, a second standard deviation, an entropy value and a second gradient norm.
[0135] It should be noted that the foregoing explanation of the embodiment of the method for evaluating the performance of the reinforcement learning algorithm based on multi-dimensional indexes is also applicable to the embodiment of the device for evaluating the performance of the reinforcement learning algorithm based on multi-dimensional indexes, which will not be described here again.
[0136] The device for evaluating the performance of the reinforcement learning algorithm based on multi-dimensional indexes according to the embodiments of the present application preprocesses the multi-dimensional index data in the reinforcement learning training process, calculates the efficiency score, the stability score and the safety score based on the preprocessed multi-dimensional index data, analyzes the correlation between the multi-dimensional indexes based on the preset Pearson correlation coefficient calculation formula according to the efficiency score, the stability score and the safety score, obtains the correlation evaluation result, and generates an optimization suggestion strategy to perform an optimization action according to the optimization suggestion strategy. In this way, the problem of single index in the existing reinforcement learning evaluation method and the difficulty in accurately reflecting the overall performance of the algorithm are solved, and more accurate and comprehensive evaluation and optimization are achieved through comprehensive analysis of multi-dimensional indexes.
[0137] Figure 4 The structure schematic diagram of the electronic device provided by the embodiments of the present application is shown. The electronic device can include:
[0138] The memory 401, the processor 402 and the computer program stored in the memory 401 and executable on the processor 402.
[0139] The processor 402 implements the method for evaluating the performance of the reinforcement learning algorithm based on multi-dimensional indexes provided in the above embodiments when executing the program.
[0140] Further, the electronic device further includes:
[0141] The communication interface 403 is configured to communicate between the memory 401 and the processor 402.
[0142] The memory 401 is configured to store a computer program executable in the processor 402.
[0143] The memory 401 can include a high-speed RAM memory, and can further include a non-volatile memory, for example, at least one disk memory.
[0144] If the memory 401, the processor 402 and the communication interface 403 are independently implemented, the communication interface 403, the memory 401 and the processor 402 can be connected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 4 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0145] Optionally, in a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are integrated on a chip, the memory 401, the processor 402 and the communication interface 403 can communicate with each other through an internal interface.
[0146] The processor 402 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0147] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the above-mentioned multi-dimensional index-based reinforcement learning algorithm performance evaluation method.
[0148] In the description of the application, reference to "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that a particular feature, structure, material, or characteristic being described is included in at least one embodiment or example of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the described specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. In addition, the usage of "N" means at least two, for example two, three or the like, unless explicitly stated otherwise.
[0149] Furthermore, the terms "first", "second", or the like, are used merely as a designation of certain elements or features of the application, and do not imply or connote relative importance or a specific order of precedence. Thus, features defined with "first", "second", etc. can include at least one of the features, either explicitly or implicitly.
[0150] Any process or method descriptions or blocks in flow charts or otherwise described herein represent embodiments of modules, segments, or portions of code which include one or more executable instructions for implementing specific logic functions or steps, and alternate implementations are possible. The preferred embodiments of this application include additional or fewer processes or methods, and the reordering of the processes or methods is possible.
[0151] The logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable instructions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a product of the manufacturing and / or processing. The computer-readable medium can include, but is not limited to, the following: an electronic connection (an electronic device with one or N wires), a portable computer diskette (a magnetic device), a RAM (random access memory), a ROM (read-only memory), an EPROM (erasable programmable ROM) or a Flash memory, an optical fiber, and a portable CD ROM. In addition, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, since the program can be electronically captured, for example, by the optically scanning the paper or other suitable medium, then electronically captured, interpreted, or processed in a suitable manner if necessary, and stored in the computer storage.
[0152] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware and in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.
[0153] Those skilled in the art can understand that all or part of the steps carried out by the above-mentioned embodiment methods can be completed by programs instructing related hardware, and the programs can be stored in a computer-readable storage medium. When the programs are executed, one or a combination of the steps of the method embodiments is included.
[0154] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can be physically present separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0155] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. A method for evaluating the performance of a reinforcement learning algorithm based on multidimensional indicators, characterized in that: The following steps are involved: Collecting multidimensional indicator data during the reinforcement learning training process, and preprocessing the multidimensional indicator data to obtain preprocessed multidimensional indicator data; Calculating an efficiency score, a stability score, and a safety score based on the preprocessed multidimensional indicator data, and analyzing the correlation between the multidimensional indicators based on the efficiency score, the stability score, and the safety score based on a preset Pearson correlation coefficient calculation formula to obtain a correlation evaluation result; An optimization suggestion strategy is generated based on the correlation evaluation result, so as to perform an optimization action according to the optimization suggestion strategy.
2. The method according to claim 1, characterized in that Generating an optimization suggestion strategy based on the correlation evaluation result, and executing an optimization action according to the optimization suggestion strategy, includes: Get the current policy entropy, current gradient update frequency, current reward signal, and current cumulative return; Based on the correlation evaluation result, the balance strategy between exploration and exploitation is adjusted according to the current policy entropy value and the current gradient update frequency, and / or the current reward function is reconstructed according to the correlation between the current reward signal and the current cumulative reward.
3. The method according to claim 1, characterized in that After calculating the efficiency score, stability score, and safety score based on the pre-processed multi-dimensional index data, the method further includes: Obtaining a comprehensive score according to the efficiency score, the stability score, and the safety score; The comprehensive score is: S total =w1S efficiency +w2S stability +w3S safety Among them, S total is the comprehensive score, W1 is the weight of the comprehensive score, S efficiency is the comprehensive score, w2 is the weight of the stability score, S stability is the stability score, w3 is the weight of the security score, S safety Score the safety.
4. The method according to claim 2, characterized in that Also includes: Based on a preset time series curve chart, the changing trend of each indicator in the multidimensional indicator data over time is monitored, and the comprehensive score is visualized.
5. The method according to claim 1, wherein The preset Pearson correlation coefficient calculation formula is: Among them, r xy is the correlation coefficient between the x variable and the y variable, n is the number of observations, and i is the i-th observation x i is the i-th x variable, is the mean of variable x, y i is the i-th y variable, is the mean of the variable y.
6. The method according to any one of claims 1 to 5, characterized in that The multidimensional indicator data includes at least one of critic-related indicators, policy-related indicators, environmental feedback data and reward signals; the critic-related indicators include at least one of the average Q value, the first standard deviation and the first gradient norm, and the policy-related indicators include at least one of the strategy mean, the second standard deviation, the entropy value and the second gradient norm.
7. A device for evaluating the performance of a reinforcement learning algorithm based on multi-dimensional indicators, characterized in that: include: An acquisition module is used to acquire multidimensional indicator data during the reinforcement learning training process and preprocess the multidimensional indicator data to obtain preprocessed multidimensional indicator data; a calculation module, configured to calculate an efficiency score, a stability score, and a safety score based on the preprocessed multidimensional indicator data, and analyze the correlation between the multidimensional indicators based on the efficiency score, the stability score, and the safety score based on a preset Pearson correlation coefficient calculation formula to obtain a correlation evaluation result; A generating module is used to generate an optimization suggestion strategy based on the correlation evaluation result, so as to perform an optimization action according to the optimization suggestion strategy.
8. The device according to claim 7, characterized in that The generating module is used to: Get the current policy entropy, current gradient update frequency, current reward signal, and current cumulative return; Based on the correlation evaluation result, the balance strategy between exploration and exploitation is adjusted according to the current policy entropy value and the current gradient update frequency, and / or the current reward function is reconstructed according to the correlation between the current reward signal and the current cumulative reward.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for evaluating the performance of a reinforcement learning algorithm based on multidimensional indicators as described in any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the reinforcement learning algorithm performance evaluation method based on multidimensional indicators as described in any one of claims 1 to 6.
Citation Information
Cited By
Mailbox security control method and device, storage medium and electronic equipment
CN121333983A