Hyper-parameter adjustment method and device based on security reinforcement learning, equipment and storage medium
By converting the optimization problem of the control strategy into a Lagrangian problem with constraints, and using gradient descent and gradient rise methods to update the parameters and Lagrangian multipliers in the iterative optimization process, the limitations of the traditional gradient descent algorithm in the optimization of complex loss function are solved, and the effective exploration and positioning of the global optimal solution in safe reinforcement learning is achieved.
Patent Information
- Application Number
- CN202510113630.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
In security reinforcement learning, traditional gradient descent algorithms are difficult to effectively explore and locate global optimal solutions for complex loss functions, especially in optimization spaces where there is a concave Pareto frontier.
The optimization problem of the control strategy is converted into a Lagrangian problem with constraints, and the main loss function, the secondary loss function and the safety threshold are introduced. The initial Lagrangian function is constructed through the Lagrangian multiplier, and the initial multiplier value is set through the dual problem. Then, during the iterative optimization process, the model parameters are updated using the gradient descent method, the gradient rise method updates the Lagrangian multiplier, and evaluates whether the strategy reaches global optimality in the enhanced Lagrangian function.
On the premise of satisfying security constraints, the main loss function is effectively minimized, which improves the ability to find global optimal solutions in complex optimization spaces, and ensures the security and efficiency of policy optimization.
Smart Images

Figure CN120046700A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of safety reinforcement learning and related technical fields. Specifically, it relates to a method, apparatus, device, and storage medium for hyperparameter tuning based on safety reinforcement learning. Background Art
[0002] Reinforcement learning (RL) is a machine learning paradigm, and its core lies in that the agent learns how to adopt the best action strategy in a specific environment by executing actions and observing environmental feedback. Safe reinforcement learning (SRL) is a branch of reinforcement learning, which particularly emphasizes incorporating safety constraints into the policy optimization process to ensure that the agent does not pose unacceptable risks during exploration and learning.
[0003] In safe reinforcement learning, some methods use an end-to-end policy network as an approximate solver for large-scale constrained optimization problems. However, when faced with loss functions composed of complex linear combinations, especially when there is a concave Pareto front in the optimization space formed by these functions, traditional gradient descent algorithms often struggle to effectively explore and locate the global optimal solution. This constitutes a major challenge in safe reinforcement learning. Summary of the Invention
[0004] Embodiments described herein provide a method, apparatus, device, and storage medium for hyperparameter tuning based on safety reinforcement learning in machine learning to solve the policy optimization problem in these application scenarios with extremely high safety requirements, and at the same time overcome the limitations of traditional gradient descent algorithms in optimizing complex loss functions.
[0005] According to a first aspect of the present disclosure, there is provided a method for hyperparameter tuning based on safety reinforcement learning, including:
[0006] Converting the optimization problem of the control policy into a Lagrangian problem with constraint conditions, specifically including: defining a primary loss function as the objective function to measure task completion, defining a secondary loss function as the safety constraint function to evaluate dangerous behaviors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process;
[0007] Introducing Lagrange multipliers, combining the primary loss function, the secondary loss function, and the safety threshold to construct an initial Lagrangian function for minimizing the primary loss function while satisfying safety constraints;
[0008] Setting an initial value for the Lagrange multiplier by solving the dual problem of the initial Lagrangian function;
[0009] Based on the initial Lagrangian function, a damping factor is introduced to construct an enhanced Lagrangian function;
[0010] Initialize the parameters of the safety reinforcement learning model to obtain an initial model configuration;
[0011] In the iterative optimization process, use the gradient descent method to update the parameters of the safety reinforcement learning model, and at the same time, use the gradient ascent method to update the Lagrange multipliers;
[0012] In each iteration, substitute the updated parameters and the updated Lagrange multipliers into the enhanced Lagrangian function for calculation, and evaluate whether the current control strategy reaches the global optimum, that is, whether the main loss function is minimized and all safety constraints are satisfied, until a global optimum solution that meets all conditions is found.
[0013] In some embodiments of the present disclosure, the step of using the gradient descent method to update the parameters of the safety reinforcement learning model includes:
[0014] Calculate the first gradient of the objective function with respect to the parameters;
[0015] Use the gradient descent method to update the parameters of the safety reinforcement learning model, and the calculation formula is: parameter = parameter - learning rate * first gradient;
[0016] The step of using the gradient ascent method to update the Lagrange multipliers includes:
[0017] Calculate the second gradient of the Lagrangian function with respect to the Lagrange multipliers;
[0018] Use the gradient ascent method to update the Lagrange multipliers, and the calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient.
[0019] In some embodiments of the present disclosure, the step of evaluating whether the current control strategy reaches the global optimum includes:
[0020] Judge whether the current control strategy realizes the minimization of the main loss function and satisfies all safety constraints;
[0021] If so, continue to iteratively update the parameters and the Lagrange multipliers;
[0022] If not, stop the iteration, consider that the global optimum solution that meets all conditions has been found, and save the updated parameters and the updated Lagrange multipliers.
[0023] In some embodiments of the present disclosure, when introducing a damping factor to construct an enhanced Lagrangian function, an adaptive adjustment strategy is adopted to dynamically adjust the value of the damping factor according to the current iteration number or the optimization progress.
[0024] In some embodiments of the present disclosure, the main loss function measures the efficiency or effect of task completion based on at least one of the following factors: time, resource consumption, accuracy, and performance metrics.
[0025] In some embodiments of the present disclosure, the secondary loss function evaluates safety based on at least one of the following potential dangerous behaviors: collision risk, out-of-operation range, and violation of operation rules.
[0026] In some embodiments of the present disclosure, the safety threshold is determined according to the safety standards, operation specifications, or empirical data of the system.
[0027] According to the second aspect of the present disclosure, a hyperparameter adjustment method based on safety reinforcement learning is applied to an autonomous driving scenario. The method includes:
[0028] Converting the optimization problem of the autonomous driving control strategy into a Lagrangian problem with constraints, specifically including: defining the main loss function as the objective function to measure the vehicle driving efficiency, defining the secondary loss function as the safety constraint function to evaluate the behavior of violating traffic rules, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process;
[0029] Introducing a Lagrange multiplier, combining the main loss function, the secondary loss function, and the safety threshold to construct an initial Lagrangian function to minimize the main loss function on the premise of satisfying the safety constraint;
[0030] By solving the dual problem of the initial Lagrangian function, setting an initial value for the Lagrange multiplier;
[0031] On the basis of the initial Lagrangian function, introducing a damping factor to construct an enhanced Lagrangian function;
[0032] Initializing the parameters of the safety reinforcement learning model to obtain an initial model configuration;
[0033] During the iterative optimization process, using the gradient descent method to update the parameters of the safety reinforcement learning model, and at the same time, using the gradient ascent method to update the Lagrange multiplier;
[0034] In each iteration, substitute the updated parameters and the updated Lagrange multipliers into the enhanced Lagrangian function for calculation, and evaluate whether the current autonomous driving control strategy meets the requirements of the highest driving efficiency and no violation of traffic rules until a global optimal solution that satisfies all conditions is found.
[0035] In some embodiments of the present disclosure, the step of updating the parameters of the safety reinforcement learning model by using the gradient descent method includes:
[0036] Calculate the first gradient of the objective function with respect to the parameters;
[0037] Update the parameters of the safety reinforcement learning model by using the gradient descent method, and the calculation formula is: parameter = parameter - learning rate * first gradient;
[0038] The step of updating the Lagrange multipliers by using the gradient ascent method includes:
[0039] Calculate the second gradient of the Lagrangian function with respect to the Lagrange multipliers;
[0040] Update the Lagrange multipliers by using the gradient ascent method, and the calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient.
[0041] In some embodiments of the present disclosure, the step of evaluating whether the current control strategy reaches the global optimum includes:
[0042] Judge whether the current control strategy realizes the minimization of the main loss function and satisfies all safety constraint conditions;
[0043] If so, continue to iteratively update the parameters and the Lagrange multipliers;
[0044] If not, stop the iteration, consider that the global optimal solution that satisfies all conditions has been found, and save the updated parameters and the updated Lagrange multipliers.
[0045] In some embodiments of the present disclosure, when constructing the enhanced Lagrangian function by introducing a damping factor, an adaptive adjustment strategy is adopted to dynamically adjust the value of the damping factor according to the current iteration number or the optimization progress.
[0046] In some embodiments of the present disclosure, the main loss function measures the efficiency or effect of task completion based on at least one of the following factors: time, resource consumption, accuracy, performance metrics.
[0047] In some embodiments of the present disclosure, the secondary loss function evaluates the safety based on at least one of the following potential dangerous behaviors: collision risk, out-of-operation range, violation of operation rules.
[0048] In some embodiments of the present disclosure, the safety threshold is determined according to the safety standards, operation specifications or empirical data of the system.
[0049] According to the third aspect of the present disclosure, a hyperparameter tuning method based on safety reinforcement learning is applied to the power grid scenario. The method includes:
[0050] Converting the optimization problem of the power grid control strategy into a Lagrangian problem with constraints, specifically including: defining the main loss function as the objective function to measure the transmission efficiency of the power grid, defining the secondary loss function as the safety constraint function to evaluate the behavior of causing a large - scale power outage risk due to operation errors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process;
[0051] Introducing Lagrange multipliers, combining the main loss function, the secondary loss function and the safety threshold to construct an initial Lagrangian function to minimize the main loss function on the premise of meeting safety constraints;
[0052] By solving the dual problem of the initial Lagrangian function, setting an initial value for the Lagrange multiplier;
[0053] On the basis of the initial Lagrangian function, introducing a damping factor to construct an enhanced Lagrangian function;
[0054] Initializing the parameters of the safety reinforcement learning model to obtain an initial model configuration;
[0055] In the iterative optimization process, using the gradient descent method to update the parameters of the safety reinforcement learning model, and at the same time, using the gradient ascent method to update the Lagrange multiplier;
[0056] In each iteration, substituting the updated parameters and the updated Lagrange multiplier into the enhanced Lagrangian function for calculation, and evaluating whether the current power grid control strategy meets the requirements of maximizing the transmission efficiency of the power grid and avoiding large - scale power outages caused by operation errors until a global optimal solution that meets all conditions is found.
[0057] In some embodiments of the present disclosure, the step of using the gradient descent method to update the parameters of the safety reinforcement learning model includes:
[0058] Calculating the first gradient of the objective function with respect to the parameters;
[0059] Using the gradient descent method to update the parameters of the safety reinforcement learning model, and the calculation formula is: parameter = parameter - learning rate * first gradient;
[0060] The step of using the gradient ascent method to update the Lagrange multiplier includes:
[0061] Calculate the second gradient of the Lagrangian function with respect to the Lagrange multiplier;
[0062] Update the Lagrange multiplier using the gradient ascent method, and the calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient.
[0063] In some embodiments of the present disclosure, the step of evaluating whether the current control strategy reaches the global optimum includes:
[0064] Determine whether the current control strategy achieves the minimization of the main loss function and satisfies all safety constraint conditions;
[0065] If so, continue to iteratively update the parameters and the Lagrange multiplier;
[0066] If not, stop the iteration, consider that the global optimum solution satisfying all conditions has been found, and save the updated parameters and the updated Lagrange multiplier.
[0067] In some embodiments of the present disclosure, when introducing a damping factor to construct an enhanced Lagrangian function, an adaptive adjustment strategy is adopted to dynamically adjust the value of the damping factor according to the current number of iterations or the optimization progress.
[0068] In some embodiments of the present disclosure, the main loss function measures the efficiency or effect of task completion based on at least one of the following factors: time, resource consumption, accuracy, performance metrics.
[0069] In some embodiments of the present disclosure, the secondary loss function evaluates safety based on at least one of the following potential dangerous behaviors: collision risk, out-of-operation range, violation of operation rules.
[0070] In some embodiments of the present disclosure, the safety threshold is determined according to the safety standards, operation specifications or empirical data of the system.
[0071] According to the fourth aspect of the present disclosure, a hyperparameter tuning method based on safety reinforcement learning is applied to a robot scenario, and the method includes:
[0072] Convert the optimization problem of the robot control strategy into a Lagrangian problem with constraint conditions, specifically including: defining the main loss function as the objective function to measure the task completion efficiency, defining the secondary loss function as the safety constraint function to evaluate the situation of damaging surrounding equipment or injuring people by mistake, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process;
[0073] Introduce Lagrange multipliers, and combine the main loss function, the secondary loss function, and the safety threshold to construct an initial Lagrangian function for minimizing the main loss function while satisfying safety constraints;
[0074] By solving the dual problem of the initial Lagrangian function, set an initial value for the Lagrange multipliers;
[0075] On the basis of the initial Lagrangian function, introduce a damping factor to construct an enhanced Lagrangian function;
[0076] Initialize the parameters of the safety reinforcement learning model to obtain an initial model configuration;
[0077] During the iterative optimization process, use the gradient descent method to update the parameters of the safety reinforcement learning model, and at the same time, use the gradient ascent method to update the Lagrange multipliers;
[0078] In each iteration, substitute the updated parameters and the updated Lagrange multipliers into the enhanced Lagrangian function for calculation, and evaluate whether the current robot control strategy meets the requirements of the highest task completion efficiency without damaging surrounding equipment or injuring people until a global optimal solution that satisfies all conditions is found.
[0079] In some embodiments of the present disclosure, the step of using the gradient descent method to update the parameters of the safety reinforcement learning model includes:
[0080] Calculate the first gradient of the objective function with respect to the parameters;
[0081] Use the gradient descent method to update the parameters of the safety reinforcement learning model, and the calculation formula is: parameter = parameter - learning rate * first gradient;
[0082] The step of using the gradient ascent method to update the Lagrange multipliers includes:
[0083] Calculate the second gradient of the Lagrangian function with respect to the Lagrange multipliers;
[0084] Use the gradient ascent method to update the Lagrange multipliers, and the calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient.
[0085] In some embodiments of the present disclosure, the step of evaluating whether the current control strategy reaches the global optimum includes:
[0086] Judge whether the current control strategy realizes the minimization of the main loss function and satisfies all safety constraint conditions;
[0087] If so, continue to iteratively update the parameters and the Lagrange multipliers;
[0088] If not, stop the iteration, consider that the globally optimal solution satisfying all conditions has been found, and save the updated parameters and the updated Lagrange multipliers.
[0089] In some embodiments of the present disclosure, when constructing an enhanced Lagrangian function by introducing a damping factor, an adaptive adjustment strategy is adopted to dynamically adjust the value of the damping factor according to the current number of iterations or the optimization progress.
[0090] In some embodiments of the present disclosure, the main loss function measures the efficiency or effect of task completion based on at least one of the following factors: time, resource consumption, accuracy, performance metrics.
[0091] In some embodiments of the present disclosure, the secondary loss function evaluates safety based on at least one of the following potential dangerous behaviors: collision risk, out-of-operation range, violation of operating rules.
[0092] In some embodiments of the present disclosure, the safety threshold is determined according to the safety standards, operating specifications or empirical data of the system.
[0093] According to the fifth aspect of the present disclosure, there is provided a hyperparameter tuning device based on safety reinforcement learning, including:
[0094] A safety constraint conversion module, configured to convert the optimization problem of the control policy into a Lagrangian problem with constraint conditions, specifically including: defining a main loss function as the objective function to measure the task completion situation, defining a secondary loss function as the safety constraint function to evaluate potential dangerous behaviors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process;
[0095] A first construction module, configured to introduce Lagrange multipliers, combine the main loss function, the secondary loss function and the safety threshold to construct an initial Lagrangian function for minimizing the main loss function under the premise of satisfying the safety constraints;
[0096] An initial value setting module, configured to set an initial value for the Lagrange multipliers by solving the dual problem of the initial Lagrangian function;
[0097] A second construction module, configured to introduce a damping factor on the basis of the initial Lagrangian function to construct an enhanced Lagrangian function;
[0098] An initialization setting module, configured to perform initialization settings on the parameters of the safety reinforcement learning model to obtain an initial model configuration:
[0099] An update module, configured to update the parameters of the safety reinforcement learning model by using the gradient descent method during the iterative optimization process, and simultaneously update the Lagrange multiplier by using the gradient ascent method;
[0100] A calculation and evaluation module, configured to substitute the updated parameters and the updated Lagrange multiplier into the augmented Lagrangian function for calculation in each iteration, and evaluate whether the current control strategy reaches the global optimum, that is, whether the main loss function is minimized and all safety constraint conditions are satisfied, until a global optimum solution that satisfies all conditions is found.
[0101] In some embodiments of the present disclosure, the update module 506 specifically includes:
[0102] A first calculation module 5061, configured to calculate a first gradient of the objective function with respect to the parameters;
[0103] A first update sub-module 5602, configured to update the parameters of the safety reinforcement learning model by using the gradient descent method, and the calculation formula is: parameter = parameter - learning rate * first gradient;
[0104] A second calculation module 5063, configured to calculate a second gradient of the Lagrangian function with respect to the Lagrange multiplier;
[0105] A first update sub-module 5604, configured to update the Lagrange multiplier by using the gradient ascent method, and the calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient.
[0106] In some embodiments of the present disclosure, the calculation and evaluation module 507 is specifically configured to:
[0107] Determine whether the current control strategy minimizes the main loss function and satisfies all safety constraint conditions;
[0108] If so, continue to iteratively update the parameters and the Lagrange multiplier;
[0109] If not, stop the iteration, consider that a global optimum solution that satisfies all conditions has been found, and save the updated parameters and the updated Lagrange multiplier.
[0110] In some embodiments of the present disclosure, the second construction module 504 is specifically configured to adopt an adaptive adjustment strategy when introducing a damping factor to construct an augmented Lagrangian function, and dynamically adjust the value of the damping factor according to the current iteration number or the optimization progress, so as to more flexibly control the stability and convergence speed of the algorithm.
[0111] In some embodiments of the present disclosure, the primary loss function measures the efficiency or effectiveness of task completion based on at least one of the following factors: time, resource consumption, accuracy, performance metrics.
[0112] In some embodiments of the present disclosure, the secondary loss function evaluates safety based on at least one of the following potential hazardous behaviors: collision risk, out-of-operation range, violation of operating rules.
[0113] In some embodiments of the present disclosure, the safety threshold is determined according to the safety standards, operating specifications or empirical data of the system.
[0114] According to a sixth aspect of the present disclosure, there is provided a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method in any one of the above embodiments are implemented.
[0115] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.
[0116] The hyperparameter tuning method and device based on secure reinforcement learning provided by the embodiments of the present disclosure first transform the control policy optimization into a constrained Lagrangian problem, define the primary and secondary loss functions and set a safety threshold; then construct an initial Lagrangian function in combination with the Lagrange multiplier to minimize the primary loss and satisfy the safety constraint; secondly, set the initial value of the multiplier by solving the dual problem, and then introduce a damping factor to construct an enhanced function; finally, after initializing the model parameters, use the gradient descent method to update the model parameters and the gradient ascent method to update the multiplier during iteration, and substitute them into the enhanced function to evaluate whether the policy is optimal, that is, the primary loss is minimized and the safety constraint is satisfied, until the global optimal solution is found. It realizes the optimization of the control policy in the processing of engineering systems or related technical fields, so as to improve the task completion efficiency or effect while ensuring safety. The hyperparameter tuning is realized by introducing a damping coefficient on the basis of the Lagrangian method, preventing the learned policy from oscillating when satisfying the constraint, and thus ensuring the smooth progress of the reinforcement learning process. The Lagrangian optimization framework has a high degree of flexibility and allows dynamic weight adjustment of the loss function. By solving the dual problem, the Lagrange multiplier can be efficiently determined, and then the loss function of the entire system can be optimized. The implemented hard constraint priority strategy ensures that the optimization process is carried out on the premise of satisfying specific key constraints. The basic differential multiplier method and its improved version (MDMM) provide an innovative strategy that can synchronously optimize the model parameters and the constraint conditions. To solve the safety problem in reinforcement learning, constraint conditions are introduced, the problem is transformed into a constrained optimization problem, and a damping coefficient is incorporated into the Lagrangian method, aiming to enhance the adjustability of the algorithm and effectively prevent the policy from oscillating when satisfying the constraint, thereby ensuring the smooth operation of the reinforcement learning process. By combining the Lagrangian optimization framework with the damping mechanism, the workload of hyperparameter tuning is significantly reduced, and the optimization effect is significantly improved.
[0117] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] In order to illustrate the technical solutions of the embodiments of the present disclosure more clearly, the drawings of the embodiments will be briefly described below. It should be understood that the following described drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure, where:
[0119] Figure 1 is a schematic flowchart of a hyperparameter tuning method based on secure reinforcement learning provided by the embodiments of the present disclosure;
[0120] Figure 2It is a schematic flowchart of a hyperparameter adjustment method based on secure reinforcement learning provided by an embodiment of the present disclosure;
[0121] Figure 3 It is a schematic flowchart of a hyperparameter adjustment method based on secure reinforcement learning provided by an embodiment of the present disclosure;
[0122] Figure 4 It is a schematic flowchart of a hyperparameter adjustment method based on secure reinforcement learning provided by an embodiment of the present disclosure;
[0123] Figure 5 It is a schematic structural diagram of a hyperparameter adjustment device based on secure reinforcement learning provided by an embodiment of the present disclosure;
[0124] Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure.
[0125] In the drawings, marks with the same last two digits correspond to the same elements. It should be noted that the elements in the drawings are schematic and not drawn to scale. Detailed Embodiments
[0126] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art without creative efforts based on the described embodiments of the present disclosure also fall within the scope of protection of the present disclosure.
[0127] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase "embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0128] The term "and / or" in this document is merely a description of the associated relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: the existence of A, the simultaneous existence of A and B, and the existence of B. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.
[0129] In addition, in all embodiments of the present disclosure, terms such as "first" and "second" are only used to distinguish one component (or a part of the component) from another component (or another part of the component).
[0130] In the description of this application, unless otherwise specified, "multiple" means two or more (including two). Similarly, "multiple groups" means two or more groups (including two groups).
[0131] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0132] Hyperparameters are some key variables used to define the model architecture and training process. Different values of them will lead to differences between models. For example, even for Convolutional Neural Networks (CNN) models, different numbers of layers will also make the models different. Hyperparameters are usually preset based on experience before the start of training, rather than learned through the training process. In deep learning, common hyperparameters include batch size, learning rate, optimizer type, learning rate decay strategy, number of iterations, number of network layers, and number of neurons in each layer, etc. The selection of these hyperparameters has a crucial impact on the performance and training efficiency of the model. To find the optimal combination of hyperparameters, it is usually necessary to adjust them through experiments and performance evaluations on the validation set to ensure that the model can converge to the optimal solution quickly and stably during the training process.
[0133] Safe Reinforcement Learning combines reinforcement learning with security-related technologies and methods to solve the problem of intelligent decision-making in security-sensitive environments. By clearly defining security constraints, introducing uncertainty handling methods, and applying core methods and principles.
[0134] Reinforcement learning is a machine learning method that mainly focuses on how to take actions in an environment to maximize a certain cumulative reward. A reinforcement learning model usually consists of an agent that learns a policy through interaction with the environment to achieve a specific goal. And what needs to be optimized is the policy, which is usually characterized as a neural network with parameters θ. In many real-world application scenarios, the agent needs to meet certain requirements in terms of security. For example, autonomous vehicles need to avoid traffic accidents, and financial systems need to prevent malicious attacks, etc. Therefore, safe reinforcement learning emerges as the times require, aiming to enable the agent to ensure the security and reliability of its decisions during the learning process.
[0135] In safe reinforcement learning, it is necessary to clearly define security constraints and incorporate them into the learning method. These constraints can include safety boundaries, security limits, and strict security rules. The agent needs to meet these security constraints during the learning process to ensure the security of its decision-making behavior.
[0136] During the reinforcement learning process, an agent usually has to face the uncertainty of the environment. Safe reinforcement learning better handles this uncertainty by introducing safety constraints and related technical methods to ensure that the agent can also make safe decisions in an uncertain environment.
[0137] The core methods of safe reinforcement learning are usually built on traditional reinforcement learning algorithms and introduce safety constraints. These algorithms learn, through continuous attempts and iterations, the strategy of maximizing the cumulative reward while satisfying the safety constraints.
[0138] Many application scenarios with extremely high safety requirements in the real world, such as autonomous driving, power grid regulation, training robots for reconnaissance or rescue, and network security, etc., involve sequential decision-making processes in uncertain environments. Most of the sequential decision-making tasks in these scenarios can be effectively handled by transforming them into reinforcement learning problems. However, in the classical reinforcement learning framework, the agent relies on the trial-and-error method for learning, and this process may bring potential safety risks. When the agent explores the unknown environment, it may choose to bear short-term losses in exchange for greater long-term benefits, which is unacceptable in scenarios with extremely high safety requirements.
[0139] Example 1
[0140] To solve the above technical problems, an embodiment of the present disclosure provides a hyperparameter tuning method based on safe reinforcement learning. Figure 1 It is a schematic flowchart of a hyperparameter tuning method based on safe reinforcement learning provided by an embodiment of the present disclosure. As Figure 1 shown, the specific process of the hyperparameter tuning method based on safe reinforcement learning includes:
[0141] S110. Convert the optimization problem of the control strategy into a Lagrangian problem with constraint conditions, specifically including: defining the main loss function as the objective function to measure the task completion (such as completion efficiency or effect), defining the secondary loss function as the safety constraint function to evaluate potential dangerous behaviors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process, thereby ensuring the safety of operations.
[0142] Optionally, the main loss function measures the efficiency or effect of task completion based on at least one of the following factors: time, resource consumption, accuracy, performance metrics.
[0143] Optionally, the secondary loss function evaluates safety based on at least one of the following potential dangerous behaviors: collision risk, out-of-operation range, violation of operating rules.
[0144] Optionally, the safety threshold is determined according to the safety standards, operation specifications or empirical data of the system.
[0145] In the specific implementation process, in combination with the safety reinforcement learning method, the main loss function L 0 (θ) is defined as the objective function, representing the performance metric to be minimized, used to measure the efficiency or effectiveness of task completion. The secondary loss function L 1 (θ) is defined as the safety constraint function, and a safety threshold ε is defined as the constraint condition of the safety constraint function to indicate the conditions that the constrained optimization needs to meet, ensuring that the value of the secondary loss function L 1 (θ) does not exceed the safety threshold ε, indicating that the constrained optimization needs to meet the condition of L 1 (θ) ≤ ε, thereby ensuring the safety of operations;
[0146] This ensures that the safety reinforcement learning model does not overly sacrifice secondary objectives when optimizing the main task.
[0147] Safety reinforcement learning has broad application prospects in multiple fields, including but not limited to the following scenarios:
[0148] 1. Autonomous driving scenario: Autonomous vehicles need to avoid traffic accidents and ensure the safety of passengers and pedestrians. Safety reinforcement learning can help autonomous vehicles learn safer driving strategies.
[0149] 2. Power grid control scenario: In the field of power grid control, it is necessary to avoid large-scale power outages caused by operation errors and ensure normal power supply. Safety reinforcement learning can ensure that the power grid does not cause large-scale power outages due to operation errors during energy transmission.
[0150] 3. Robot control scenario: In the field of robot control, safety reinforcement learning can ensure that robots do not damage surrounding equipment or harm humans when performing tasks.
[0151] 4. Cybersecurity scenario: In the field of cybersecurity, safety reinforcement learning can be used to optimize firewall policies, malware detection policies, etc., improving the efficiency and accuracy of cybersecurity systems.
[0152] Only some application scenarios are listed above as examples, but the present invention is not limited to these situations. Any other scenario applying the safety reinforcement learning of the embodiments of the present disclosure, as long as it involves a sequential decision-making process in an uncertain environment and needs to consider the interaction safety issues between the agent and the surrounding environment, is within the protection scope of the present invention.
[0153] The present invention aims to provide a safer and more effective strategy optimization method to adapt to various complex application scenarios.
[0154] In a specific embodiment, as an example, in an autonomous driving scenario, the primary loss function includes: driving time and energy consumption, and the secondary loss function includes: the number and severity of traffic rule violations.
[0155] In a specific embodiment, as another example, in a power grid control scenario, the primary loss function includes: the efficiency of maximum energy transmission and the number of power grid dispatching operations, and the secondary loss function includes: large-scale power outages caused by operation errors.
[0156] In a specific embodiment, as another example, in a robot control scenario such as training for reconnaissance or rescue, the primary loss function includes: task completion, and the secondary loss function includes: damaging expensive equipment of one's own side or not accidentally injuring innocent people.
[0157] In a specific embodiment, as another example, in a network security control scenario, the primary loss function includes: network information transmission efficiency, and the secondary loss function includes: no firewall being invaded by malware.
[0158] S120. Introduce Lagrange multipliers, combine the primary loss function, the secondary loss function, and the security threshold to construct an initial Lagrangian function for minimizing the primary loss function while satisfying security constraints, thereby maximizing task efficiency.
[0159] In a specific implementation manner, introduce Lagrange multiplier λ, and combine the primary loss function L 0 (θ), the secondary loss function L 1 (θ), and the security threshold ε to construct an initial Lagrangian function. The formula is:
[0160] L(θ, λ) = L 0 (θ) - λ(ε - L 1 (θ)); (1)
[0161] Where θ is the model parameter, L 0 (θ) is the primary loss function, L 1 (θ) is the secondary loss function, ε is the security threshold, and λ is the Lagrange multiplier. That is, the secondary loss function is added as a constraint condition. Among them, the Lagrange multiplier λ is a non-negative multiplier; the initial Lagrangian function transforms the optimization problem into an unconstrained optimization problem through introducing the Lagrange multiplier, which is convenient for solving.
[0162] S130. By solving the dual problem of the initial Lagrangian function, set an initial value for the Lagrange multiplier to provide a basis for the subsequent optimization iteration process.
[0163] Before starting the optimization iteration, it is necessary to determine the initial value of the Lagrange multiplier by solving the dual problem, which is the key to starting the optimization process.
[0164] In the specific implementation process, solve the dual problem of the Lagrangian function to find the Lagrange multiplier that maximizes the dual function;
[0165] Use numerical optimization methods to solve the dual problem to obtain an initial estimate of the Lagrange multiplier.
[0166] S140. Based on the initial Lagrangian function, introduce a damping factor to construct an enhanced Lagrangian function to enhance the degree of satisfaction of the constraint conditions.
[0167] In the specific implementation process, introduce a damping factor ζ and construct an enhanced Lagrangian function based on the initial Lagrangian function. The formula is:
[0168] L(θ,λ)=L 0 (θ)-(λ - ζ)(ε - L 1 (θ)); (2)
[0169] Among them, θ is the parameter of the safety reinforcement learning model, L 0 (θ) is the main loss function, L 1 (θ) is the secondary loss function, ε is the safety threshold, λ is the Lagrange multiplier, ζ is the damping factor, and the damping factor is determined by formula (3) as follows:
[0170] ζ=ξ(ε - L 1 (θ)); (3)
[0171] Among them, ξ is the learning rate of the damping factor, θ is the parameter of the safety reinforcement learning model, ε is the safety threshold, and L 1 (θ) is the secondary loss function. The purpose of introducing the damping factor is to prevent oscillations when satisfying the constraints, control the change of the policy during the update process, and prevent over-adjustment.
[0172] In the specific implementation, when introducing the damping factor to construct the enhanced Lagrangian function, an adaptive adjustment strategy is adopted to dynamically adjust the value of the damping factor according to the current iteration number or optimization progress to more flexibly control the stability and convergence speed of the algorithm.
[0173] The enhanced Lagrangian function may have stronger constraint capabilities, which helps the model better meet the safety requirements to enhance the degree of satisfaction of the constraint conditions.
[0174] S150. Initialize the parameters of the safety reinforcement learning model to obtain an initial model configuration.
[0175] In the specific implementation process, the parameters θ of the safety reinforcement learning model are initialized to obtain a preliminary model configuration, so as to obtain an initial model configuration and provide a starting point for subsequent iterative optimization.
[0176] S160. In the iterative optimization process, the gradient descent method is used to update the parameters of the safety reinforcement learning model to further improve the efficiency or effect of task execution. At the same time, the gradient ascent method is used to update the Lagrange multiplier to strengthen the influence of safety constraints and ensure that the safety requirements are always met during the optimization process.
[0177] Optionally, the gradient descent method is at least one of the following gradient descent algorithms: stochastic gradient descent, batch gradient descent, mini-batch gradient descent.
[0178] Optionally, the gradient ascent method is at least one of the following gradient ascent algorithms: stochastic gradient ascent, batch gradient ascent, mini-batch gradient ascent.
[0179] In a specific implementation manner, the gradient descent method is used to update the parameters θ of the safety reinforcement learning model to minimize the Lagrangian function, so as to further improve the efficiency or effect of task execution.
[0180] Step 1: Calculate the first gradient of the main loss function with respect to the model parameters.
[0181] Step 2: Use the gradient descent method to update the model parameters. The calculation formula is: model parameters = model parameters - learning rate * first gradient, that is, θ = θ + learning rate * first gradient.
[0182] In a specific implementation manner, the gradient ascent method is used to update the Lagrange multiplier λ to strengthen the influence of safety constraints and ensure that the safety requirements are always met during the optimization process, so as to adjust the balance between the main loss function and this loss function.
[0183] Step 1: Calculate the second gradient of the Lagrangian function with respect to the Lagrange multiplier.
[0184] Step 2: Use the gradient ascent method to update the Lagrange multiplier. The calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient, that is, λ = λ + learning rate * second gradient.
[0185] S170. In each iteration, the updated parameters and the updated Lagrange multiplier are substituted into the augmented Lagrangian function for calculation to evaluate whether the current control strategy reaches the global optimum, that is, whether the main loss function is minimized and all safety constraint conditions are met, until a global optimum solution that meets all conditions is found.
[0186] Safety policy iteration ensures that the agent's decisions gradually approach the optimal policy while satisfying safety constraints by continuously updating the policy.
[0187] In the specific implementation process, the step of evaluating whether the current control policy reaches the global optimum includes:
[0188] Determine whether the current control policy achieves the minimization of the main loss function and satisfies all safety constraint conditions;
[0189] If so, continue to iteratively update the parameters and the Lagrange multipliers;
[0190] If not, stop the iteration, consider that the global optimum solution satisfying all conditions has been found, and save the updated parameters and the updated Lagrange multipliers.
[0191] The hyperparameter tuning method based on safe reinforcement learning provided by the embodiments of the present disclosure first transforms the control policy optimization into a constrained Lagrangian problem; defines the main and secondary loss functions and sets the safety threshold; secondly, constructs an initial Lagrangian function in combination with the Lagrange multipliers to minimize the main loss and satisfy the safety constraints; then sets the initial value of the multiplier by solving the dual problem, and introduces a damping factor to construct an enhanced function; finally, after initializing the model parameters, uses the gradient descent method to update the model parameters and the gradient ascent method to update the multipliers in the iteration, and substitutes them into the enhanced function to evaluate whether the policy is optimal, that is, the main loss is minimized and the safety constraints are satisfied, until the global optimum solution is found. By introducing a damping coefficient on the basis of the Lagrangian method, the adjustability of the algorithm is improved, and the oscillation of the learned policy when satisfying the constraints is prevented, thereby ensuring the smooth progress of the reinforcement learning process.
[0192] Embodiment 2
[0193] Figure 2 It is a flowchart of a hyperparameter tuning method based on safe reinforcement learning provided by the embodiments of the present disclosure. As Figure 2 shown, the hyperparameter tuning method based on safe reinforcement learning is applied to the field of autonomous driving vehicles. The specific process includes:
[0194] A reinforcement learning problem of an autonomous driving vehicle. In the optimization of the control policy of an autonomous driving vehicle, it is necessary to improve the driving efficiency (such as reaching the destination in the shortest time) on the premise of ensuring driving safety (avoiding dangerous behaviors such as collisions), and it is necessary to optimize the policy to maximize the driving efficiency (main loss function), while ensuring compliance with traffic rules (secondary loss function).
[0195] S210. Convert the optimization problem of the autonomous driving control strategy into a Lagrangian problem with constraints, specifically including: setting the goal of the autonomous driving vehicle as optimizing the driving strategy to maximize the driving efficiency (defined as the main loss function, such as driving time or energy consumption), while strictly complying with traffic rules (defined as the secondary loss function, such as the number or severity of traffic rule violations). To ensure safety, the value of the secondary loss function needs to be limited to not exceed a certain set safety threshold.
[0196] S220. Introduce Lagrange multipliers, combine the main loss function, the secondary loss function, and the safety threshold to construct an initial Lagrangian function, which is used to minimize the loss of driving efficiency (i.e., the main loss function) as much as possible on the premise of complying with traffic rules (i.e., the constraint conditions of the secondary loss function), so as to maximize the driving efficiency.
[0197] S230. By solving the dual problem of the initial Lagrangian function, set the initial value for the Lagrange multipliers, providing a basis for the subsequent optimization iteration process;
[0198] In addition, to ensure that critical safety rules are not violated under any circumstances, a hard constraint priority strategy is implemented. The hard constraint priority strategy is to ensure that critical traffic safety rules are not violated under any circumstances.
[0199] S240. On the basis of the initial Lagrangian function, introduce a damping factor to construct an enhanced Lagrangian function;
[0200] The purpose of introducing the damping mechanism is to control the change of the strategy during the update process and prevent over-adjustment.
[0201] S250. Initialize the parameters of the safety reinforcement learning model to obtain an initial model configuration, providing a starting point for subsequent iterative optimization;
[0202] S260. During the iterative optimization process, use the gradient descent method to update the parameters of the safety reinforcement learning model to further improve the efficiency or effect of task execution. At the same time, use the gradient ascent method to update the Lagrange multipliers to strengthen the influence of safety constraints. To prevent the strategy from over-adjusting during the update process, a damping mechanism is introduced to control the change of the strategy.
[0203] S270. In each iteration, substitute the updated parameters and the updated Lagrange multipliers into the enhanced Lagrangian function for calculation, and evaluate whether the current autonomous driving control strategy reaches the global optimum, that is, whether it meets the requirements of the highest efficiency and no violation of traffic rules, until a global optimum solution that meets all conditions is found.
[0204] Through the above embodiments, an efficient and safe autonomous driving strategy is achieved, while meeting the performance and constraint requirements in reinforcement learning.
[0205] For Embodiment 2, since it corresponds to Embodiment 1 and is a common application scenario of Embodiment 1, the relevant parts can be referred to the partial description of the method embodiment.
[0206] Embodiment 3
[0207] Figure 3 is a schematic flowchart of a hyperparameter tuning method based on safe reinforcement learning provided by an embodiment of the present disclosure. As Figure 3 shown, the hyperparameter tuning method based on safe reinforcement learning is applied to the power grid control field. The specific process includes:
[0208] For a reinforcement learning problem of power grid control, it is necessary to optimize the strategy to maximize the efficiency of energy transmission and reduce the number of power grid scheduling operations (main loss function), while ensuring to avoid large-scale power outages caused by operation errors (secondary loss function).
[0209] S310. Convert the optimization problem of the power grid control strategy into a Lagrangian problem with constraint conditions, specifically including: setting the goal of power grid control as optimizing the energy transmission strategy to maximize the transmission efficiency (defined as the main loss function, such as the efficiency of energy transmission and the number of power grid scheduling operations), while strictly complying with the operation rules (defined as the secondary loss function, such as the risk of large-scale power outages caused by operation errors). To ensure safety, the value of the secondary loss function needs to be limited within a certain set safety threshold to ensure the safety of power grid operation.
[0210] S320. Introduce Lagrange multipliers, combine the main loss function, the secondary loss function, and the safety threshold to construct an initial Lagrangian function, so as to minimize the transmission efficiency loss on the premise of meeting the operation rules and maximize the transmission efficiency;
[0211] S330. By solving the dual problem of the initial Lagrangian function, set an initial value for the Lagrange multiplier to provide a basis for the subsequent optimization iteration process;
[0212] In addition, to ensure no operation errors in any case, a hard constraint priority strategy is implemented. The hard constraint priority strategy is to ensure no operation errors in any case.
[0213] S340. On the basis of the initial Lagrangian function, introduce a damping factor to construct an enhanced Lagrangian function;
[0214] The purpose of introducing the damping mechanism is to control the change of the strategy during the update process and prevent over-adjustment.
[0215] S350. Initialize the parameters of the safety reinforcement learning model to obtain an initial model configuration, providing a starting point for subsequent iterative optimization.
[0216] S360. During the iterative optimization process, use the gradient descent method to update the parameters of the safety reinforcement learning model to further improve the efficiency or effectiveness of task execution. At the same time, use the gradient ascent method to update the Lagrange multiplier to strengthen the influence of safety constraints and ensure that safety requirements are always met during the optimization process.
[0217] S370. In each iteration, substitute the updated parameters and the updated Lagrange multiplier into the augmented Lagrangian function for calculation to evaluate whether the current power grid control strategy reaches the global optimum, that is, whether it maximizes the transmission efficiency of the power grid and effectively avoids the requirement of large-scale power outages caused by operation errors, until a global optimum solution that meets all conditions is found.
[0218] Through the above embodiments, an efficient power grid driving strategy is realized, while meeting the performance and constraint requirements in reinforcement learning.
[0219] For Embodiment 3, since it corresponds to Embodiment 1 and is a common specific application scenario of Embodiment 1, the relevant parts can be referred to the partial description of the method embodiment.
[0220] Embodiment 4
[0221] Figure 4 is a flowchart of a hyperparameter tuning method based on safety reinforcement learning provided by an embodiment of the present disclosure. As Figure 4 shown, the hyperparameter tuning method based on safety reinforcement learning is applied to the field of robot control.
[0222] Reinforcement learning technology is applied to train robots of types such as reconnaissance or rescue to perform specific tasks. However, in training or actual combat applications, these robots are strictly prohibited from attempting to damage high-value equipment of their own side or accidentally injure irrelevant personnel. Preventing the occurrence of such dangerous behaviors is even more important than the completion of the task itself.
[0223] To address the safety issues in reinforcement learning, constraint conditions are introduced, thus transforming the problem into a constrained optimization problem. A damping coefficient is incorporated into the Lagrangian method, aiming to improve the adjustability of the algorithm and effectively prevent the policy from oscillating when meeting the constraint conditions, thereby ensuring the smooth progress of the reinforcement learning process.
[0224] A reinforcement learning problem for training robots of types such as reconnaissance or rescue requires optimizing the policy to complete the task (main loss function), while ensuring not to damage surrounding equipment or accidentally injure people (secondary loss function).
[0225] S410. Convert the optimization problem of the robot control strategy into a Lagrangian problem with constraints, specifically including: setting the goal of robot control as optimizing the control strategy to maximize the task completion efficiency (defined as the main loss function, such as the task completion efficiency), while not damaging surrounding equipment or hurting people by mistake (defined as the secondary loss function, such as the number or severity of violations of traffic rules). To ensure safety, the value of the secondary loss function needs to be limited to not exceed a certain set safety threshold.
[0226] S420. Introduce Lagrange multipliers, combine the main loss function, the secondary loss function and the safety threshold to construct an initial Lagrangian function, so as to minimize the task efficiency loss on the premise of meeting safety rules and maximize the task completion efficiency;
[0227] S430. By solving the dual problem of the initial Lagrangian function, set the initial value for the Lagrange multiplier to provide a basis for the subsequent optimization iteration process;
[0228] In addition, to ensure that key safety rules are not violated under any circumstances, a hard constraint priority strategy is implemented. The hard constraint priority strategy is to ensure the safety of operations under any circumstances.
[0229] S440. On the basis of the initial Lagrangian function, introduce a damping factor to construct an enhanced Lagrangian function;
[0230] The purpose of introducing the damping mechanism is to control the change of the control strategy during the update process and prevent over-adjustment.
[0231] S450. Initialize the parameters of the safety reinforcement learning model to obtain an initial model configuration and provide a starting point for subsequent iterative optimization;
[0232] S460. During the iterative optimization process, use the gradient descent method to update the parameters of the safety reinforcement learning model to further improve the efficiency or effect of task execution. At the same time, use the gradient ascent method to update the Lagrange multiplier to strengthen the influence of safety constraints and ensure that safety requirements are always met during the optimization process;
[0233] S470. In each iteration, substitute the updated parameters and the updated Lagrange multiplier into the enhanced Lagrangian function for calculation to evaluate whether the current robot control strategy reaches the global optimum, that is, whether the requirements of the highest efficiency and no damage to surrounding equipment or no hurting people by mistake are met, until a global optimum solution that meets all conditions is found.
[0234] Through the above embodiments, a strategy for training robots of types such as reconnaissance or rescue that is both efficient and safe is realized, while meeting the performance and constraint requirements in reinforcement learning.
[0235] For Embodiment 4, since it corresponds to Embodiment 1 and is a common application scenario of Embodiment 1, the relevant parts can be referred to the partial description of the method embodiment.
[0236] Only some application scenarios are listed above as examples, but the present invention is not limited to these situations. Any other scenario applying the safety reinforcement learning of the embodiments of the present disclosure, as long as it involves a sequential decision-making process in an uncertain environment and needs to consider the interaction safety problem between the agent and the surrounding environment, is within the protection scope of the present invention. The present invention aims to provide a safer and more effective strategy optimization method to adapt to various complex application scenarios.
[0237] Embodiment 5
[0238] Based on the above embodiments, the embodiments of the present disclosure further provide a hyperparameter adjustment device based on safety reinforcement learning, as Figure 5 shown. The hyperparameter adjustment device based on safety reinforcement learning includes:
[0239] A safety constraint conversion module 501, configured to convert the optimization problem of the control policy into a Lagrangian problem with constraint conditions, specifically including: defining a main loss function as the objective function to measure the task completion situation, defining a secondary loss function as the safety constraint function to evaluate potential dangerous behaviors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process;
[0240] A first construction module 502, configured to introduce a Lagrange multiplier, combine the main loss function, the secondary loss function, and the safety threshold to construct an initial Lagrangian function to minimize the main loss function on the premise of satisfying the safety constraint;
[0241] An initial value setting module 503, configured to set an initial value for the Lagrange multiplier by solving the dual problem of the initial Lagrangian function;
[0242] A second construction module 504, configured to introduce a damping factor on the basis of the initial Lagrangian function to construct an enhanced Lagrangian function;
[0243] An initialization setting module 505, configured to perform initialization settings on the parameters of the safety reinforcement learning model to obtain an initial model configuration:
[0244] An update module 506, configured to update the parameters of the safety reinforcement learning model by using the gradient descent method during the iterative optimization process, and at the same time, update the Lagrange multiplier by using the gradient ascent method;
[0245] A calculation and evaluation module 507, which is used to substitute the updated parameters and the updated Lagrange multipliers into the enhanced Lagrange function for calculation in each iteration, evaluate whether the current control strategy reaches the global optimum, that is, whether the main loss function is minimized and all safety constraint conditions are satisfied, until a global optimum solution that meets all conditions is found.
[0246] In a specific implementation manner, the update module 506 specifically includes:
[0247] A first calculation module 5061, which is used to calculate the first gradient of the objective function with respect to the parameters;
[0248] A first update sub-module 5062, which is used to update the parameters of the safety reinforcement learning model by using the gradient descent method, and the calculation formula is: parameter = parameter - learning rate * first gradient;
[0249] A second calculation module 5063, which is used to calculate the second gradient of the Lagrange function with respect to the Lagrange multipliers;
[0250] A first update sub-module 5064, which is used to update the Lagrange multipliers by using the gradient ascent method, and the calculation formula is: Lagrange multiplier = Lagrange multiplier + learning rate * second gradient.
[0251] In a specific implementation manner, the calculation and evaluation module 507 is specifically used for:
[0252] Judge whether the current control strategy minimizes the main loss function and satisfies all safety constraint conditions;
[0253] If so, continue to iteratively update the parameters and the Lagrange multipliers;
[0254] If not, stop the iteration, consider that the global optimum solution that meets all conditions has been found, and save the updated parameters and the updated Lagrange multipliers.
[0255] In a specific implementation manner, the second construction module 504 is specifically used to adopt an adaptive adjustment strategy when introducing a damping factor to construct an enhanced Lagrange function, and dynamically adjust the value of the damping factor according to the current iteration number or the optimization progress, so as to more flexibly control the stability and convergence speed of the algorithm.
[0256] Optionally, the main loss function measures the efficiency or effect of task completion based on at least one of the following factors: time, resource consumption, accuracy, performance metrics.
[0257] Optionally, the secondary loss function evaluates safety based on at least one of the following potential dangerous behaviors: collision risk, out-of-operation range, violation of operating rules.
[0258] Optionally, the safety threshold is determined according to the safety standards, operation specifications or empirical data of the system.
[0259] The hyperparameter tuning device based on safety reinforcement learning provided by the embodiments of the present disclosure realizes hyperparameter tuning by introducing a damping coefficient on the basis of the Lagrangian method, preventing the oscillation of the learned policy when satisfying the constraints, and thus ensuring the smooth progress of the reinforcement learning process.
[0260] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.
[0261] The embodiments of the present application also provide a computer device. Specifically, please refer to Figure 6 , Figure 6 , which is the basic structural block diagram of the computer device in this embodiment.
[0262] The computer device includes a memory 610 and a processor 620 that communicate with each other through a system bus. It should be noted that only the computer device with components 610-620 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0263] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server and other computing devices. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad or a voice control device and other means.
[0264] The memory 610 includes at least one type of readable storage medium, which includes non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. The RAM may include static RAM or dynamic RAM. In some embodiments, the memory 610 may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 610 may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device. Of course, the memory 610 may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory 610 is generally used to store the operating system and various application software installed on the computer device, such as the program code of the above method. In addition, the memory 610 may also be used to temporarily store various types of data that have been output or will be output.
[0265] The processor 620 is generally used to execute the overall operations of the computer device. In this embodiment, the memory 610 is used to store program code or instructions, and the program code includes computer operation instructions. The processor 620 is used to execute the program code or instructions stored in the memory 610 or process data, such as running the program code of the above method.
[0266] In this text, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus system can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0267] Another embodiment of the present application further provides a computer-readable medium, which can be a computer-readable signal medium or a computer-readable medium. A processor in the computer reads the computer-readable program code stored in the computer-readable medium, so that the processor can execute the functional actions specified in each step or the combination of steps in the above method; and generate a device for implementing the functional actions specified in each block or the combination of blocks in the block diagram.
[0268] The computer-readable medium includes but is not limited to electronic, magnetic, optical, electromagnetic, infrared memories or semiconductor systems, devices or apparatuses, or any suitable combination of the foregoing. The memory is used to store program code or instructions, and the program code includes computer operation instructions. The processor is used to execute the program code or instructions of the above method stored in the memory.
[0269] For the settings of the memory and the processor, reference can be made to the description of the foregoing computer device embodiments, and details are not described herein again.
[0270] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0271] In each embodiment of the present application, each functional unit or module can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0272] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0273] Unless otherwise explicitly stated in the context, the singular forms of the words used in this specification and the appended claims include the plural, and vice versa. Thus, when referring to the singular, the plural of the corresponding term is usually included. Similarly, the terms "comprising" and "including" will be interpreted as inclusive rather than exclusive. Likewise, the term "including" and "or" should be interpreted as inclusive, unless such an interpretation is explicitly prohibited in this specification. Where the term "example" is used in this specification, especially when it is located after a group of terms, the "example" is merely exemplary and illustrative and should not be considered exclusive or extensive.
[0274] Further aspects and scopes of adaptability become apparent from the description provided herein. It should be understood that the various aspects of this application can be implemented alone or in combination with one or more other aspects. It should also be understood that the description herein and the specific embodiments are for illustrative purposes only and are not intended to limit the scope of this application.
[0275] The above has described several embodiments of the present disclosure in detail. However, obviously, those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The protection scope of the present disclosure is defined by the appended claims.
Claims
1. A hyperparameter adjustment method based on secure reinforcement learning, characterized in that: include: The optimization problem of the control strategy is converted into a Lagrangian problem with constraints, including: defining a primary loss function as an objective function to measure the task completion, defining a secondary loss function as a safety constraint function to evaluate dangerous behaviors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process; Introducing Lagrangian multipliers, combining the main loss function, the secondary loss function and the safety threshold, and constructing an initial Lagrangian function to minimize the main loss function under the premise of satisfying safety constraints; Setting an initial value for the Lagrangian multiplier by solving the dual problem of the initial Lagrangian function; Based on the initial Lagrangian function, a damping factor is introduced to construct an enhanced Lagrangian function; Initialize the parameters of the security reinforcement learning model to obtain the initial model configuration; In the iterative optimization process, the parameters of the security reinforcement learning model are updated using the gradient descent method, and at the same time, the Lagrange multiplier is updated using the gradient ascent method; In each iteration, the updated parameters and the updated Lagrangian multipliers are substituted into the enhanced Lagrangian function for calculation to evaluate whether the current control strategy reaches the global optimum, that is, whether the main loss function is minimized and all safety constraints are met, until the global optimal solution that meets all conditions is found.
2. The method according to claim 1, characterized in that The step of updating the parameters of the security reinforcement learning model using the gradient descent method includes: Calculating a first gradient of the objective function with respect to a parameter; The parameters of the security reinforcement learning model are updated using the gradient descent method, and the calculation formula is: parameter = parameter - learning rate * first gradient; The step of updating the Lagrange multiplier using the gradient ascent method comprises: Calculating a second gradient of the Lagrangian function with respect to the Lagrangian multiplier; The Lagrange multiplier is updated using the gradient ascent method, and the calculation formula is: Lagrange multiplier=Lagrange multiplier+learning rate*second gradient.
3. The method according to claim 1, characterized in that The step of evaluating whether the current control strategy reaches the global optimum includes: Determine whether the current control strategy achieves the minimization of the main loss function and satisfies all safety constraints; If yes, continue to iteratively update the parameters and the Lagrange multiplier; If not, the iteration is stopped, it is considered that a global optimal solution satisfying all conditions has been found, and the updated parameters and the updated Lagrange multipliers are saved.
4. The method according to claim 1, characterized in that: When introducing the damping factor to construct the enhanced Lagrangian function, an adaptive adjustment strategy is adopted to dynamically adjust the value of the damping factor according to the current number of iterations or optimization progress.
5. A hyperparameter adjustment method based on secure reinforcement learning, characterized in that: Applied to the autonomous driving scenario, the method includes: The optimization problem of the autonomous driving control strategy is converted into a Lagrangian problem with constraints, including: defining a primary loss function as an objective function to measure the vehicle's driving efficiency, defining a secondary loss function as a safety constraint function to evaluate violations of traffic rules, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process; Introducing Lagrangian multipliers, combining the main loss function, the secondary loss function and the safety threshold, and constructing an initial Lagrangian function to minimize the main loss function under the premise of satisfying safety constraints; Setting an initial value for the Lagrangian multiplier by solving the dual problem of the initial Lagrangian function; Based on the initial Lagrangian function, a damping factor is introduced to construct an enhanced Lagrangian function; Initialize the parameters of the security reinforcement learning model to obtain the initial model configuration; In the iterative optimization process, the parameters of the security reinforcement learning model are updated using the gradient descent method, and at the same time, the Lagrange multiplier is updated using the gradient ascent method; In each iteration, the updated parameters and the updated Lagrangian multipliers are substituted into the enhanced Lagrangian function for calculation to evaluate whether the current autonomous driving control strategy achieves the highest driving efficiency and does not violate traffic regulations, until a global optimal solution that meets all conditions is found.
6. A hyperparameter adjustment method based on secure reinforcement learning, characterized in that: Applied to the power grid scenario, the method includes: The optimization problem of the power grid control strategy is converted into a Lagrangian problem with constraints, including: defining a primary loss function as an objective function to measure the transmission efficiency of the power grid, defining a secondary loss function as a safety constraint function to evaluate the behavior of large-scale power outage risks caused by operational errors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process; Introducing Lagrangian multipliers, combining the main loss function, the secondary loss function and the safety threshold, and constructing an initial Lagrangian function to minimize the main loss function under the premise of satisfying safety constraints; Setting an initial value for the Lagrangian multiplier by solving the dual problem of the initial Lagrangian function; Based on the initial Lagrangian function, a damping factor is introduced to construct an enhanced Lagrangian function; Initialize the parameters of the security reinforcement learning model to obtain the initial model configuration; In the iterative optimization process, the parameters of the security reinforcement learning model are updated using the gradient descent method, and at the same time, the Lagrange multiplier is updated using the gradient ascent method; In each iteration, the updated parameters and the updated Lagrangian multipliers are substituted into the enhanced Lagrangian function for calculation to evaluate whether the current power grid control strategy achieves the requirements of maximizing the transmission efficiency of the power grid and avoiding large-scale power outages caused by operational errors, until a global optimal solution that meets all conditions is found.
7. A hyperparameter adjustment method based on secure reinforcement learning, characterized in that: Applied to the robot scenario, the method includes: The optimization problem of the robot control strategy is converted into a Lagrangian problem with constraints, including: defining a main loss function as an objective function to measure the efficiency of task completion, defining a secondary loss function as a safety constraint function to evaluate the damage to surrounding equipment or accidental injury to people, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process; Introducing Lagrangian multipliers, combining the main loss function, the secondary loss function and the safety threshold, and constructing an initial Lagrangian function to minimize the main loss function under the premise of satisfying safety constraints; Setting an initial value for the Lagrangian multiplier by solving the dual problem of the initial Lagrangian function; Based on the initial Lagrangian function, a damping factor is introduced to construct an enhanced Lagrangian function; Initialize the parameters of the security reinforcement learning model to obtain the initial model configuration; In the iterative optimization process, the parameters of the security reinforcement learning model are updated using the gradient descent method, and at the same time, the Lagrange multiplier is updated using the gradient ascent method; In each iteration, the updated parameters and the updated Lagrangian multipliers are substituted into the enhanced Lagrangian function for calculation to evaluate whether the current robot control strategy achieves the highest task completion efficiency without damaging surrounding equipment or accidentally injuring people, until a global optimal solution that meets all conditions is found.
8. A hyperparameter adjustment device based on secure reinforcement learning, characterized in that: include: A safety constraint conversion module is used to convert the optimization problem of the control strategy into a Lagrangian problem with constraints, specifically including: defining a main loss function as an objective function to measure the task completion, defining a secondary loss function as a safety constraint function to evaluate potential dangerous behaviors, and setting a safety threshold to ensure that the value of the secondary loss function does not exceed the safety threshold during the optimization process; A first construction module is used to introduce a Lagrangian multiplier, combine the main loss function, the secondary loss function and the safety threshold, and construct an initial Lagrangian function to minimize the main loss function under the premise of satisfying the safety constraint; An initial value setting module, used for setting an initial value for the Lagrangian multiplier by solving the dual problem of the initial Lagrangian function; A second construction module is used to introduce a damping factor on the basis of the initial Lagrangian function to construct an enhanced Lagrangian function; The initialization setting module is used to initialize the parameters of the security reinforcement learning model to obtain the initial model configuration: An updating module, used to update the parameters of the security reinforcement learning model using a gradient descent method during an iterative optimization process, and at the same time, update the Lagrange multiplier using a gradient ascent method; The calculation and evaluation module is used to substitute the updated parameters and the updated Lagrangian multipliers into the enhanced Lagrangian function for calculation in each iteration to evaluate whether the current control strategy reaches the global optimum, that is, whether the main loss function is minimized and all safety constraints are met, until the global optimal solution that meets all conditions is found.
9. A computer device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Umbrella ladder system modeling optimization method and system based on mechanics and parameter optimization
CN120449510A
Aircraft recovery scheduling method and equipment based on adaptive Lagrange multiplier
CN120725393A
DAB converter control method, system and equipment based on current stress optimization
CN121012313A
Obstacle avoidance control method based on dynamic obstacle trajectory prediction
CN121596879A