Layered adaptive safety reinforcement learning system and method applied to patrol robot

By using a hierarchical adaptive safety reinforcement learning system combined with a safety factor genetic algorithm to optimize the assessment of potential hazards, the problem of insufficient safety and efficiency of patrol robots in complex environments is solved, achieving intelligent switching and the best balance between safety and efficiency.

CN121300346APending Publication Date: 2026-01-09THE FIRST RES INST OF MIN OF PUBLIC SECURITY +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511378000.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing reinforcement learning methods struggle to ensure safety and reliability in complex and dynamic environments for patrol robot tasks. They cannot effectively distinguish between different hazard levels and semantic-level safety rules, and lack adaptability and flexibility, resulting in insufficient strategies for robots when dealing with emergency avoidance and efficient patrolling.

Method used

A hierarchical adaptive safety reinforcement learning system is adopted, which includes a bottom-level safety factor policy pool and a high-level adaptive safety meta-policy. The potential risk assessment function is optimized through a safety factor genetic algorithm, and a safety factor policy pool is formed by combining multiple safety factor reward functions. The low-level policy is intelligently switched in different situations through the adaptive safety meta-policy.

Benefits of technology

It significantly improves the safety and mission efficiency of patrol robots, enabling them to intelligently switch safety strategies in dynamic environments, avoid dangerous behaviors, achieve the best balance between safety and efficiency, and improve adaptability and mission success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005613326770000051
    Figure BDA0005613326770000051
  • Figure BDA0005613326770000111
    Figure BDA0005613326770000111
  • Figure FDA0005613326760000021
    Figure FDA0005613326760000021
Patent Text Reader

Abstract

The invention discloses a hierarchical self-adaptive safety reinforcement learning system and method applied to a patrol robot, and the method comprises the steps: introducing a hierarchical safety factor strategy pool and a self-adaptive safety element strategy, and optimizing a potential risk assessment function through an SFGA; according to the method, the safety violation times (such as entering a forbidden area and colliding with a dynamic / static object too close to the forbidden area) can be comprehensively and remarkably reduced, and the average minimum safety distance is increased, so that dangerous behaviors are effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics, and more specifically to a hierarchical adaptive safety reinforcement learning system and method for patrol robots. Background Technology

[0002] With the development of artificial intelligence and robotics, reinforcement learning has shown great potential in autonomous navigation and task execution for robots. Patrol robots, as an important application, have broad prospects in security, surveillance, and other fields. However, traditional reinforcement learning methods, while pursuing task efficiency (such as fastest exploration and highest score), often struggle to guarantee safety and reliability in complex and dynamic environments. The working environment of patrol robots is typically a dynamic and partially observable physical space, including static objects (such as buildings and walls), dynamic objects (such as pedestrians and vehicles), and special areas (such as patrol waypoints, restricted areas, and charging stations). Robots need to receive high-dimensional information from the environment (such as their own position, velocity, sensor data, map information, and task progress) to make decisions. Their action space typically includes navigation actions (such as setting linear and angular velocities) and task-related actions (such as controlling the camera and activating warning lights). In reinforcement learning, the design of the reward function is crucial; it defines the task objective and undesirable behaviors, typically including task rewards (such as reaching the patrol point or completing the patrol) and penalties (such as collisions, entering restricted areas, and inefficient behavior).

[0003] Shariq Iqbal et al. from UCLA proposed a framework for designing intrinsic rewards that take into account areas already explored by other agents, thereby encouraging cooperative behavior among agents. Furthermore, they developed a method to learn how to dynamically choose among multiple exploration strategies (or "exploration modalities") to maximize the final task reward (extrinsic reward). Specifically, this method employs a hierarchical policy. A high-level "meta-controller" is responsible for selecting from a set of low-level policies trained based on different intrinsic rewards. These low-level controllers, guided by specific intrinsic rewards, are then responsible for learning the specific action strategies for all agents.

[0004] Currently, reinforcement learning is mainly applied in two ways to robot patrol tasks:

[0005] 1. Efficiency-driven approach: Minimize trajectory length or maximize coverage, performs well in exploration and map building, and usually only includes a simple, uniform collision penalty as a safety constraint.

[0006] 2. Behavior-driven approach: Learning different exploration "behaviors" through intrinsic rewards aims to solve the problem of exploration efficiency under sparse rewards.

[0007] However, existing methods have their limitations:

[0008] 1. Limitations of efficiency-driven approaches: These approaches fall short when dealing with complex safety requirements. They often fail to differentiate between different levels of danger (e.g., "near a wall" versus "near a moving child"), and cannot handle semantic-level safety rules (such as "no entry" zones). A single collision penalty is insufficient to address the diverse safety risks in the real world.

[0009] 2. Limitations of Behavior-Driven Approaches: Although methods such as Coordinated Exploration via Intrinsic Rewards for Multi-Agent Reinforcement Learning (CE) introduce hierarchical frameworks and intrinsic rewards, these intrinsic rewards are designed for exploration efficiency, not safety. Directly applying them to patrol tasks may allow robots to learn efficient exploration, but the safety of the patrol process cannot be guaranteed. For example, a robot might enter restricted areas or approach pedestrians in order to "explore" new areas.

[0010] 3. Lack of safety diversity and adaptability: Existing methods typically employ a fixed, single safety penalty function. This makes it difficult for robots to intelligently switch safety strategies in different scenarios, such as "emergency pedestrian avoidance" and "efficient patrolling in open areas." In dynamically changing environments, this "one-size-fits-all" safety strategy lacks flexibility. It may lead to robots being unable to make optimal decisions when pedestrian avoidance is prioritized due to excessive concern about hitting walls, or constantly worrying about all potential dangers, resulting in extremely low task execution efficiency or even "retreating." Summary of the Invention

[0011] To address the shortcomings of existing technologies, this invention aims to provide a hierarchical adaptive safety reinforcement learning system and method for use in patrol robots.

[0012] To achieve the above objectives, the present invention adopts the following technical solution:

[0013] A hierarchical adaptive safety reinforcement learning system for patrol robots includes a bottom-level safety factor policy pool and a high-level adaptive safety meta-policy.

[0014] The security factor policy pool stores a set of low-level policies driven by different security factor reward functions and with different security preferences. The goal of each low-level policy is to maximize a composite reward. The composite reward includes the main patrol task reward and the intrinsic security factor reward. The main patrol task reward is an extrinsic reward shared by all low-level policies, which refers to the positive reward obtained when the patrol robot reaches a predetermined patrol point. The intrinsic security factor reward is a set of rewards corresponding to different security dimensions, used to drive the low-level policies to learn different security preferences.

[0015] Each low-level policy is trained using a standard reinforcement learning algorithm, with the goal of maximizing the composite reward consisting of the main patrol task reward and the intrinsic reward of the specified security factor; after training, the low-level policies form a security factor policy pool.

[0016] The adaptive security meta-policy is a high-level controller used to select the most suitable low-level policy from the underlying security factor policy pool for execution in each decision cycle. The input of the adaptive security meta-policy is global or local environmental information, and the goal of the adaptive security meta-policy is to learn a selection function to maximize the long-term reward of the main patrol mission.

[0017] Furthermore, the low-level strategies include restricted area avoidance strategy, static obstacle avoidance strategy, dynamic obstacle avoidance strategy, efficiency-first static obstacle avoidance strategy, and smooth action strategy; the intrinsic reward of the safety factor for each low-level strategy is specifically as follows:

[0018] Restricted Area Avoidance Strategy: When a patrol robot enters a pre-set restricted area, it will receive a huge negative reward;

[0019] Static obstacle avoidance strategy: The reward is inversely proportional to the distance between the patrol robot and the nearest static obstacle; the closer the distance, the greater the negative reward.

[0020] Dynamic obstacle avoidance strategy: The reward is inversely proportional to the distance between the patrol robot and the nearest dynamic individual, and the penalty coefficient is much higher than that of static obstacle avoidance;

[0021] Smooth motion strategy: Negative rewards are given for the patrol robot’s sharp turns, rapid acceleration and deceleration to encourage smooth and predictable movements;

[0022] Efficiency-first static obstacle avoidance strategy: This strategy enables patrol robots to prioritize task efficiency while avoiding static obstacles, allowing them to patrol at higher speeds in open areas.

[0023] Furthermore, the adaptive safety meta-policy is trained using the policy gradient method, and the reward it receives is the extrinsic reward obtained from actually performing the task, namely the reward for the main patrol task. Through learning, the adaptive safety meta-policy can discover and understand that following different low-level policies in different situations is necessary to better complete the final task.

[0024] Furthermore, the safety factor reward function is shown in the following formula:

[0025]

[0026] Where, λ danger The intensity of the penalty for danger is represented by a safety factor genetic algorithm that evolves; D thresh is the hazard activation threshold, s represents the current state information perceived by the patrol robot from the environment at a certain point in time, a represents the action performed by the patrol robot in state s, and s' represents the new state that the patrol robot transitions to after performing action a in state s; II(*) is the indicator function, which is 1 when condition * is true, and 0 otherwise; D(s,a) represents the potential hazard assessment function.

[0027] The relevant features representing the state-action pairs of the patrol robot are defined as phi_i(s,a). In the safety factor genetic algorithm, the chromosome encoded by the gene is a real-number vector containing the weights (w_1,...,w_N) of each feature and the intensity λ of the danger penalty. danger N is the total number of features;

[0028] Using each chromosome generated by the safety factor genetic algorithm as a set of parameters, perform the following steps:

[0029] A1. Use this set of parameters to construct the safety factor reward function R. SF ;

[0030] A2. Select a standard reinforcement learning algorithm to train the low-level policy or the adaptive safety meta-policy;

[0031] A3. Use the total reward R_total = r_task + R SF The patrol robot is trained for M episodes for evaluation; the trained patrol robot is evaluated, and the fitness function F is calculated; the fitness function F aims to maximize task utility while minimizing actual constraint violations and prediction errors.

[0032] Furthermore, phi_i(s,a) includes:

[0033] phi_1 represents the reciprocal of the Euclidean distance between the patrol robot and the nearest static obstacle after the robot performs action 'a'; a larger value indicates a closer distance and a higher level of danger.

[0034] phi_2 represents the reciprocal of the Euclidean distance between the patrol robot and the nearest dynamic individual after the robot performs action 'a'; a larger value indicates a closer distance and a higher level of danger.

[0035] phi_3 indicates whether the patrol robot has entered a known restricted area after performing action a. If it has entered the restricted area, the value is 1; otherwise, it is 0.

[0036] phi_4 represents the normalized magnitude of the patrol robot's current linear velocity;

[0037] phi_5 represents the normalized magnitude of the patrol robot's current angular velocity;

[0038] phi_6 represents the reciprocal of the distance between the patrol robot and the least recently visited patrol point.

[0039] Furthermore, define task utility J util Define the average task utility return (e.g., number of patrol points completed or efficiency) for the patrol robot during the evaluation phase; define the actual constraint violation J. viol_actual Let J_danger_prediction_accuracy be the actual average constraint violation cost of the patrol robot during the evaluation phase; the accuracy of hazard prediction is defined as the accuracy of the prediction of D(s,a), specifically by calculating the activation of D(s,a), i.e., D(s,a) exceeds D_s, where D_s,a, is the activation of D(s,a). thresh However, the false positive rate when no actual violation occurred, and the false negative rate when an actual violation occurred but D(s,a) was not activated;

[0040] The fitness formula is: F = W util *J util -W viol *J viol_actual -W pre_error *(FalsePositiveRate+FalseNegativeRate)

[0041] Wherein, fitness F is the final score used in the genetic algorithm to evaluate the quality of each group's safety factors, and is used to maximize the F value; FalsePositiveRate represents the false alarm rate, which is the proportion of the potential hazard assessment function D(s,a) being activated, i.e., D(s,a) exceeding the threshold, but no violation actually occurring; FalseNegativeRate represents the false negative rate, which is the proportion of violations actually occurring, but D(s,a) not being activated; weight W util W viol and W pre_error These are all hyperparameters that need to be adjusted through training to balance the three objectives of task completion efficiency, safety efficiency, and prediction accuracy.

[0042] Furthermore, the genetic operators of the safety factor genetic algorithm employ crossover with standard real-number encoding.

[0043] This invention also provides a method for patrolling robots to operate using the above-mentioned hierarchical adaptive security reinforcement learning system, the specific process of which is as follows:

[0044] The patrol robot perceives current status information from the environment;

[0045] The high-level adaptive safety meta-policy receives the current state information perceived by the patrol robot as input, and then selects the most suitable low-level policy from the underlying safety factor policy pool based on the current environmental state. The selected low-level policy outputs a navigation action based on its learned behavioral preferences. The patrol robot executes the navigation action output by the low-level policy, interacts with the environment, the environmental state changes, and an immediate reward r_task and a potential safety factor penalty are generated. The total reward signal R_total, obtained by summing the immediate reward r_task and the potential safety factor penalty, is R_total = R_task + R SF This is further used to train an adaptive security meta-policy, enabling it to learn when to choose which security preferences while maximizing rewards for major patrol missions in the long run.

[0046] The beneficial effects of this invention are as follows:

[0047] 1. Significantly improve safety: By introducing a hierarchical safety factor policy pool and an adaptive safety meta-policy, and by optimizing the potential hazard assessment function through SFGA (Safety Factor Genetic Algorithm), this invention can comprehensively and significantly reduce the number of safety violations (such as entering restricted areas or colliding too closely with dynamic / static objects) and increase the average minimum safe distance, thereby effectively avoiding dangerous behaviors.

[0048] 2. Achieving the optimal balance between safety and efficiency: This invention can intelligently switch between "emergency pedestrian avoidance" and "efficient patrolling in open areas" based on dynamic environmental changes, avoiding the overly conservative approach of a single, fixed safety strategy. This maximizes the efficiency of patrol missions while ensuring safety. It is expected to approach the efficiency-first model in terms of mission efficiency metrics, while significantly exceeding existing methods in terms of safety metrics.

[0049] 3. Enhance the adaptive capabilities of intelligent agents: The adaptive safety meta-strategy enables patrol robots to understand and adapt to different safety contexts, exhibiting more intelligent decision-making capabilities that are closer to human judgment, such as slowing down and carefully avoiding pedestrians, while accelerating in open areas.

[0050] 4. Providing more refined and proactive safety guidance: SFGA can learn to identify and penalize complex, non-obvious hazard patterns through evolutionary weights and penalty intensity. This enables patrol robots to anticipate and avoid potential dangers, rather than passively penalizing collisions that have already occurred. This makes safety assurance more proactive and detailed.

[0051] 5. Enhanced interpretability: By analyzing the feature weights evolved from SFGA, we can gain insights into which environmental factors are identified as hazardous factors by the system, providing guidance for the design of safety strategies.

[0052] 6. Improve mission success rate and stability: By balancing safety and efficiency, patrol robots can complete complex patrol tasks more stably, reducing mission failures caused by safety violations or excessive conservatism. Detailed Implementation

[0053] The present invention will be further described below. It should be noted that this embodiment is based on the present technical solution and provides detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.

[0054] Example 1

[0055] This embodiment provides a hierarchical adaptive safety reinforcement learning system for patrol robots, including a bottom-level safety factor policy pool and a high-level adaptive safety meta-policy.

[0056] The Safety Factor Policy Pool stores a set of low-level policies driven by different safety factor reward functions and with different safety preferences. The goal of each low-level policy is to maximize a composite reward. The composite reward includes the main patrol task reward (r_task) and the intrinsic safety reward.

[0057] The primary patrol mission reward is an external reward shared by all lower-level strategies, which refers to the positive reward obtained when the patrol robot reaches a predetermined patrol point.

[0058] The intrinsic reward of the security factor is a set of rewards corresponding to different security dimensions, used to drive low-level policies to learn different security preferences; for example:

[0059] In the restricted area avoidance strategy (r_forbidden), when a patrol robot enters a pre-defined restricted area, it receives a huge negative reward.

[0060] In the static obstacle avoidance strategy (r_static), the reward is inversely proportional to the distance between the patrol robot and the nearest static obstacle (wall, furniture); the closer the distance, the greater the negative reward.

[0061] In the dynamic obstacle avoidance strategy (r_dynamic), the reward is inversely proportional to the distance between the patrol robot and the nearest dynamic individual (simulated pedestrian), and the penalty coefficient is much higher than that of the static obstacle avoidance strategy.

[0062] In the smooth action strategy (r_action), negative rewards are given to the patrol robot’s actions such as sharp turns, rapid acceleration / deceleration (i.e., large acceleration or angular velocity) to encourage its smooth and predictable actions.

[0063] In the efficiency-first static obstacle avoidance strategy (r_static_eff), the reward is similar to that of the static obstacle avoidance strategy, but the penalty coefficient is lower. The purpose is to make the patrol robot focus more on task efficiency while avoiding static obstacles, allowing it to patrol at a higher speed in open areas.

[0064] Each low-level policy (pi_i) is trained using a standard reinforcement learning algorithm (such as PPO, SAC, etc.) with the goal of maximizing a composite reward consisting of the primary patrol task reward and the intrinsic reward of a specified security factor. For example, one low-level policy might focus on maximizing (r_task + r_forbidden), while another might focus on maximizing (r_task + r_dynamic). After training, these low-level policies form a security factor policy pool.

[0065] The Adaptive Safety Meta-Policy is a high-level controller that does not directly select actions. Instead, it selects the most suitable low-level policy from the underlying safety factor policy pool at each decision cycle (or at regular intervals). The input to the Adaptive Safety Meta-Policy can be global or local environmental information, such as whether the patrol robot's sensors have detected a dynamic individual nearby, whether the patrol robot is approaching a restricted area, and the patrol robot's current speed. The goal of the Adaptive Safety Meta-Policy is to learn a selection function (m_safety|state) to maximize the long-term reward of the primary patrol task (r_task).

[0066] Specifically, in this embodiment, the adaptive security meta-policy is trained using a policy gradient method, and its reward is the extrinsic reward obtained from actually performing patrol tasks, i.e., the primary patrol task reward (r_task). Through learning, the adaptive security meta-policy can discover and understand that following different low-level policies in different situations is necessary to better complete the final task.

[0067] For example:

[0068] In open areas, the adaptive security meta-policy will select a low-level policy that prioritizes efficiency (such as r_static or r_action) to patrol smoothly at a faster speed.

[0069] When the sensor detects a pedestrian nearby, the adaptive safety meta-policy will immediately switch to a lower-level policy (such as r_dynamic) that focuses on dynamic obstacle avoidance, prioritizing a safe distance from the pedestrian, even if this sacrifices patrol efficiency.

[0070] When approaching the boundary of the restricted area, the adaptive safety meta-policy will select a low-level policy that focuses on avoiding the restricted area (such as r_forbidden) to ensure that the boundary is not crossed.

[0071] In this embodiment, the definition and weight of the safety factor are changed from fixed preset values ​​to automatic discovery and optimization by a genetic algorithm.

[0072] The safety factor reward function is shown in the following formula:

[0073]

[0074] Where, λ danger The intensity of the penalty for danger is represented by a safety factor genetic algorithm that evolves; D thresh Here, s represents the hazard activation threshold, s represents the current state information perceived by the patrol robot from the environment at a certain point in time, a represents the action performed by the patrol robot in state s, and s' represents the new state transitioned to by the patrol robot after performing action a in state s. II(*) is the indicator function, which is 1 when condition * is true, and 0 otherwise; D(s,a) represents the potential hazard assessment function.

[0075] This embodiment introduces a safety factor genetic algorithm (SFGA) to optimize the potential hazard assessment function D(s,a) and its penalty mechanism, thereby achieving more refined and proactive safety guidance for patrol robots.

[0076] The relevant features representing the state-action pairs of the patrol robot are defined as phi_i(s,a), including:

[0077] phi_1 represents the reciprocal of the Euclidean distance between the patrol robot and the nearest static obstacle after the robot performs action 'a'. A larger value indicates a closer distance and a higher level of danger.

[0078] phi_2 represents the reciprocal of the Euclidean distance between the patrol robot and the nearest dynamic individual (pedestrian) after the robot performs action 'a'. A larger value indicates a closer distance and a higher level of danger. The penalty weight w_2 for feature phi_2 is expected to be larger than the penalty weight w_1 for phi_1.

[0079] phi_3 indicates whether the patrol robot has entered a known restricted area after performing action 'a' (0 or 1). The value is 1 if it enters the restricted area, and 0 otherwise.

[0080] phi_4 represents the normalized magnitude of the patrol robot's current linear velocity. High speed itself is not dangerous, but when approaching a boundary or when there are pedestrians, its penalty weight w_4 may significantly affect the overall hazard assessment D(s,a).

[0081] phi_5 represents the normalized magnitude of the patrol robot's current angular velocity. Similar to phi_4, its contribution to the hazard level is reflected through learned weights and combinations with other features.

[0082] phi_6 represents the reciprocal of the distance between the patrol robot and the least recently visited patrol point. This helps SFGA balance security and mission efficiency.

[0083] In the aforementioned safety factor genetic algorithm, the chromosome encoded by the gene is a real-number vector containing the weights (w_1,...,w_N) of each feature and the intensity λ of the danger penalty. danger N is the total number of features. For example, for the 6 features defined above, the chromosomes would be (w_1, w_2, w_3, w_4, w_5, w_6, λ). danger If D thresh As an evolutionary parameter, it is also included in the chromosome.

[0084] For each chromosome (i.e., a set of parameters) generated by the safety factor genetic algorithm, perform the following steps:

[0085] A1. Use this set of parameters to construct the safety factor reward function R. SF .

[0086] A2. Select a standard reinforcement learning algorithm (RL algorithm, such as PPO or SAC) to train the low-level policy or the adaptive security meta-policy. In this embodiment, optimization can be performed on a specific low-level policy (e.g., a single policy that integrates all security factors), or the policy selection mechanism of the adaptive security meta-policy can be optimized.

[0087] A3. Use the total reward R_total = r_task + R SF The patrol robot is trained for M episodes (e.g., 100 episodes) for evaluation. The trained patrol robot is evaluated, and its fitness function F is calculated. The fitness function F aims to maximize task utility while minimizing actual constraint violations and prediction errors.

[0088] Define task utility J utilDefine the average task utility return (e.g., number of patrol points completed or efficiency) for the patrol robot during the evaluation phase; define the actual constraint violation J. viol_actual The actual average constraint violation rate / cost of the patrol robot during the evaluation phase (e.g., the number of times it enters restricted areas, the number of times it is less than the safety threshold when it is within a certain distance from a dynamic pedestrian); the danger prediction accuracy (J_danger_prediction_accuracy) is defined as the accuracy of the prediction of D(s,a), specifically by calculating the activation of D(s,a) (exceeding D...). thresh However, the false positive rate when no actual violation occurred, and the false negative rate when an actual violation occurred but D(s,a) was not activated.

[0089] The fitness formula is: F = W util *J util -W viol *J viol_actual -W pre_error *(FalsePositiveRate+FalseNegativeRate)

[0090] Wherein, fitness F is the final score used in the genetic algorithm to evaluate the quality of each group's safety factors, and is used to maximize the F value; J util That is, the average return during the evaluation of patrol robots; J viol_actual This refers to the cost of assessing violations by the patrol robot during the evaluation process; FalsePositiveRate represents the false alarm rate, which is the proportion of the potential hazard assessment function D(s,a) being activated, i.e., D(s,a) exceeding the threshold, but no violation actually occurring; FalseNegativeRate represents the false negative rate, which is the proportion of violations that actually occurred, but D(s,a) was not activated. Weight W util W viol and W pre_error These are all hyperparameters that need to be adjusted through training to balance the three objectives of task completion efficiency, safety efficiency, and prediction accuracy. For example, W viol It can be set very high to punish any security violation.

[0091] In this embodiment, the genetic operators of the safety factor genetic algorithm adopt standard real number encoded crossover (such as BLX-alpha) and mutation (such as Gaussian perturbation).

[0092] Example 2

[0093] This embodiment provides a method for a patrol robot to operate using a hierarchical adaptive security reinforcement learning system. The specific process is as follows:

[0094] The patrol robot perceives current status information from the environment (including its own status, environmental perception information, map information, and task-related information).

[0095] The high-level adaptive safety meta-policy receives the current state information perceived by the patrol robot as input, and then selects the most suitable low-level policy from the lower-level safety factor policy pool based on the current environmental state. The selected low-level policy outputs navigation actions (such as linear velocity and angular velocity) based on its learned behavioral preferences. The patrol robot executes the navigation actions output by the low-level policy, interacts with the environment, the environmental state transitions, and an immediate reward (r_task) and a potential safety factor penalty (if activated) are generated. The total reward signal (R_total = r_task + R) is obtained by summing the immediate reward (r_task) and the potential safety factor penalty. SF This is further used to train an adaptive security meta-policy, enabling it to learn when to choose which security preferences while maximizing rewards for major patrol missions in the long run.

[0096] Example 3

[0097] This embodiment aims to provide simulation examples of Embodiments 1 and 2.

[0098] I. Experimental Environment and Task Setup

[0099] (1) Basic environment: A large-scale, complex indoor simulation environment similar to MARVEL is adopted, specifically a 90m×90m area.

[0100] (2) Safety elements:

[0101] (2.1) Static restricted areas: Clearly define areas on the map where patrol robots are absolutely prohibited from entering, such as hazardous material storage areas and sensitive equipment areas.

[0102] (2.2) Dynamic pedestrians: Several simulated pedestrians are set up in the environment. They move along preset (such as moving along a corridor) or random paths to simulate the flow of people in the real environment.

[0103] (2.3) Static obstacles: existing walls, pillars, desks, furniture, etc. in the environment, which are fixed objects that the patrol robot needs to avoid.

[0104] (2.4) Patrol Task: Set 10-15 preset patrol key points. The goal of the multi-agent team (e.g., consisting of 2-4 patrol robots) is to collaboratively visit all these points in the shortest time and with the least total trajectory length.

[0105] (3) Patrol Robot Hardware and Simulation:

[0106] (3.1) Patrol robot type: simulate a patrol robot with differential drive capability, such as a cleaning robot or logistics robot chassis.

[0107] (3.2) Sensors: Equipped with LiDAR for acquiring point cloud information and distance of surrounding objects, and a camera for visual perception, and can obtain its own precise position and orientation through SLAM algorithm.

[0108] (3.3) Simulation platform: Use physical simulation engines such as Gazebo or Unity to build an environment to ensure the authenticity of physical dynamics and sensor data.

[0109] II. Construction and Training of the Underlying Security Factor Policy Pool

[0110] 1) Number of strategies: Five low-level strategies are constructed, each driven by a specific safety factor reward function. These include a no-go zone avoidance strategy, a static obstacle avoidance strategy, an efficiency-first static obstacle avoidance strategy, a dynamic obstacle avoidance strategy, and a smooth action strategy.

[0111] 1.1) Forbidden Zone Avoidance Strategy (r_forbidden):

[0112] Reward function: R = r_task + r_forbidden. Where r_forbidden is a huge negative reward (e.g., -1000), triggered when the patrol robot's position (x, y) enters a predefined restricted area.

[0113] Training objective: To enable patrol robots to learn to detour around restricted areas even when the shortest path passes through them, in order to avoid entering them.

[0114] 1.2) Static obstacle avoidance strategy (r_static):

[0115] Reward function: R = r_task + r_static. Here, r_static is inversely proportional to the distance d_static between the patrol robot and the nearest static obstacle. For example, r_static = -k_1 / d_static, where k_1 is a scalar representing a "penalty coefficient," used to adjust the magnitude of the penalty the robot receives for approaching a static obstacle. The penalty intensity increases when d_static is less than a preset threshold (e.g., 0.5 meters).

[0116] Training objective: To enable patrol robots to learn to maintain a safe distance from static obstacles such as walls and furniture during patrols.

[0117] 1.3) Efficiency-first static obstacle avoidance strategy (r_static_eff):

[0118] Reward function: R = r_task + r_static_eff. The efficiency-first static obstacle avoidance strategy has a lower penalty coefficient for r_static_eff than r_static, for example, r_static_eff = -0.5*k_1 / d_static. This allows it to prioritize task efficiency while avoiding obstacles, maintaining higher speeds in open areas.

[0119] Training objective: To prioritize improving patrol efficiency while ensuring basic static obstacle avoidance.

[0120] 1.4) Dynamic obstacle avoidance strategy (r_dynamic):

[0121] Reward function: R = r_task + r_dynamic. r_dynamic is inversely proportional to the distance d_dynamic between the patrol robot and the nearest dynamic individual (pedestrian). The penalty coefficient k_2 is much higher than static obstacle avoidance; for example, r_dynamic = -k_2 / d_dynamic. It is triggered when d_dynamic is less than a safety threshold (e.g., 1.5 meters) and a significant negative reward (e.g., -10 per step) is given.

[0122] Training objective: To enable patrol robots to learn to significantly slow down, detour, or stop and wait in areas with pedestrians, prioritizing pedestrian safety.

[0123] 1.5) Smooth Action Strategy (r_action):

[0124] Reward function: R = r_task + r_action. Here, r_action provides a negative reward for the sum of the squares of the patrol robot's acceleration and angular velocity, for example, r_action = -k_3*(accel2 + angular_vel2), encouraging smooth and predictable movements. Here, accel2 represents the square of the linear acceleration, which measures the drastic change in the patrol robot's forward or backward speed; while angular_vel2 represents the square of the angular velocity, which measures the drastic change in the robot's turning speed.

[0125] Training objective: To enable patrol robots to learn smooth and predictable movement patterns, reducing sudden stops and turns.

[0126] 2) Training process of low-level strategies:

[0127] Reinforcement learning algorithm: All low-level policies are trained using the Proximal Policy Optimization (PPO) algorithm. PPO is a policy gradient algorithm suitable for continuous action spaces.

[0128] Network structure: The network structure of each low-level policy is a multilayer perceptron (MLP). The input layer receives the processed environmental state (such as obstacle distance, pedestrian relative position, and self-velocity), and the output layer is the mean and variance, which are used for Gaussian policy sampling.

[0129] Training duration: Each low-level policy is trained independently for approximately 5*10^6 hours. 6 Each low-level policy used to construct the security factor policy pool needs to interact with, learn, and optimize the environment for approximately five million timesteps in the simulation environment to ensure that the policy fully converges and stably grasps its specific security preferences.

[0130] III. Construction and Training of High-Level Adaptive Security Meta-Strategy

[0131] 3.1) Input: The input to the adaptive security meta-policy is a vector containing various environmental information;

[0132] For example: the distance to the nearest dynamic individual and its relative velocity (detected and tracked via camera and LiDAR); the distance to the nearest restricted area boundary; the patrol robot's current linear and angular velocities; mission progress (number of patrol points visited); and an indicator showing whether there are any nearby unvisited patrol points.

[0133] 3.2) Reinforcement learning algorithm: The adaptive security element policy is also trained using the PPO algorithm. Its action space is discrete, which corresponds to selecting a low-level policy from the security factor policy pool (e.g., selecting r_forbidden, r_static, r_dynamic, etc.).

[0134] 3.3) Rewards: The reward function of the adaptive security meta-policy is directly derived from the primary patrol task reward (r_task), which is the reward obtained when the patrol robot completes a patrol point visit. The adaptive security meta-policy learns when to switch low-level policies by maximizing the long-term primary patrol task reward.

[0135] 3.4) The training process of the adaptive security meta-policy is as follows:

[0136] Network Structure: The network structure of the adaptive security meta-policy is also MLP.

[0137] Training duration: After the low-level policies are trained, the adaptive safety meta-policy undergoes approximately 106 time steps of training. During training, the adaptive safety meta-policy interacts with the environment, observes the impact of different low-level policies on task completion in different contexts, and thus learns the optimal policy selection strategy.

[0138] IV. Performance Indicators:

[0139] 4.1) Task efficiency index: Task completion time / total trajectory length: the lower the better.

[0140] 4.2) Mission success rate: The probability of completing all patrol points within the specified time; the higher the better.

[0141] 4.3) Safety indicators:

[0142] 4.3.1) Total number of safety violations; the total number of safety violations includes:

[0143] 4.3.1.1) Number of times entering the restricted area.

[0144] 4.3.1.2) The number of times the dynamic pedestrian distance is less than the safety threshold (e.g., 1.5 meters).

[0145] 4.3.1.3) The number of times the distance to a static obstacle is less than a safe threshold (e.g., 0.3 meters).

[0146] The total number of safety violations is the most direct evidence for measuring safety.

[0147] 4.3.2) Average minimum safe distance: The average minimum distance between the patrol robot and dynamic / static objects throughout the entire mission. The higher the better.

[0148] 4.3.3) Motion smoothness: the average curvature or acceleration of the trajectory, the lower the better.

[0149] V. Expected Outcomes

[0150] This embodiment will comprehensively and significantly outperform all baseline methods (efficiency-first model, single fixed security policy, CE exploration framework) in terms of security metrics.

[0151] In terms of task efficiency metrics, this embodiment far exceeds the baseline of the "single fixed security policy" (because it is not bound by overly conservative security rules) and approaches the level of the "efficiency-first model," achieving the best balance between security and efficiency.

[0152] This embodiment successfully demonstrates intelligent behavior: when near pedestrians, the robot slows down and carefully avoids them (the meta-policy is r_dynamic), while in open areas it accelerates forward (the meta-policy is other policies, such as r_static_eff or r_action).

[0153] Example 4

[0154] Phase 1: SFGA Operation and Security Factor Function Generation

[0155] Objective: To run SFGA and evolve a series of candidate security factor reward function parameters.

[0156] Environment: The same environment used to train the patrol robot in Example 3.

[0157] SFGA settings:

[0158] Population size: for example, 50-100 chromosomes.

[0159] Number of iterations: for example, 100-200 generations, or until the fitness converges.

[0160] Crossover rate and mutation rate: Set according to standard GA practices, for example, crossover rate 0.8 and mutation rate 0.1.

[0161] Patrol robot training time (for fitness assessment): For each chromosome assessment, the patrol robot is trained for 10,000-100,000 time steps to ensure it has enough time to learn, but not too long to cause the overall SFGA running time to be too long.

[0162] Output: After SFGA finishes running, select the top K chromosomes in fitness (e.g., K=5), and their corresponding parameters are the generated safety factor reward function series.

[0163] Phase Two: Verification and Comparison of Safety Factor Functions

[0164] Objective: To verify the effectiveness of the security factor reward function generated by SFGA and compare it with benchmark algorithms.

[0165] Environment: The same patrol robot patrol simulation environment as in Phase 1 is used, and it can be extended to more complex variants (such as increasing pedestrian density, introducing random faults, etc.).

[0166] Selected safety factor reward function: Select several representative functions from the series of outputs from Phase 1 (e.g., those with the highest fitness and different trade-offs between safety and utility).

[0167] Patrol robot: Uses the same standard RL algorithm (PPO) as in SFGA fitness evaluation.

[0168] Training and Assessment:

[0169] For each selected safety factor reward function, use R_total = r_task + R SF Training patrol robots.

[0170] Simultaneously train the benchmark algorithm.

[0171] All training sessions were run for a sufficient number of steps to observe the final performance and convergence.

[0172] Each algorithm / configuration runs multiple random seeds (e.g., 5-10), recording the mean and confidence interval to ensure statistical significance of the results.

[0173] Evaluation metrics: The same task efficiency and safety metrics as in Example 3 were used.

[0174] Key metrics: average actual constraint violation rate per episode, average task utility return per episode, and security-weighted utility (SWU) score.

[0175] Auxiliary / analysis metrics: activation frequency of D(s,a), accuracy of D(s,a) predictions (how well they match actual dangerous situations), learning curve (how rewards and constraint violations change with the number of training steps), and stability of the training process.

[0176] The benchmark algorithm includes:

[0177] ① Standard RL (no safety): For example, PPO, which does not use any safety factors or constraint treatment.

[0178] ② Single fixed security strategy: All security reward factors (r_forbidden, r_static, r_dynamic, r_action) are added together with fixed weights to form a single, fixed composite reward function to train the robot, similar to the "non-adaptive oracle" in the CE paper.

[0179] ③Safe RL based on Lagrange multipliers: For example, PPO-Lagrangian or SAC-Lagrangian, which transform the constraints into optimization problems by introducing Lagrange multipliers.

[0180] ④CE Exploration Framework: The best-performing exploration strategies from the CE paper are directly applied to patrol missions without adding any additional safety bonuses.

[0181] To further verify the effectiveness of the hazard assessment:

[0182] Recording and Analysis: During the evaluation phase, when D(s,a) exceeds D thresh Record the current state (s, a) and the trajectory of the subsequent steps. Analyze whether these state-action pairs marked as "dangerous" actually pose a high risk of causing an actual violation.

[0183] Visualization: In the simulation environment, the value or activation state of D(s,a) can be visualized on the map to observe whether it matches dangerous areas such as restricted areas, pedestrian positions, and static obstacles.

[0184] For those skilled in the art, various corresponding changes and modifications can be made based on the above technical solutions and concepts, and all such changes and modifications should be included within the protection scope of the claims of this invention.

Claims

1. A hierarchical adaptive safety reinforcement learning system for patrol robots, characterized in that, This includes the underlying security factor policy pool and the high-level adaptive security meta-policy; The security factor policy pool stores a set of low-level policies driven by different security factor reward functions and with different security preferences. The goal of each low-level policy is to maximize a composite reward. The composite reward includes the main patrol task reward and the intrinsic security factor reward. The main patrol task reward is an extrinsic reward shared by all low-level policies, which refers to the positive reward obtained when the patrol robot reaches a predetermined patrol point. The intrinsic security factor reward is a set of rewards corresponding to different security dimensions, used to drive the low-level policies to learn different security preferences. Each low-level policy is trained using a standard reinforcement learning algorithm, with the goal of maximizing the composite reward consisting of the main patrol task reward and the intrinsic reward of the specified security factor; after training, the low-level policies form a security factor policy pool. The adaptive security meta-policy is a high-level controller used to select the most suitable low-level policy from the underlying security factor policy pool for execution in each decision cycle. The input of the adaptive security meta-policy is global or local environmental information, and the goal of the adaptive security meta-policy is to learn a selection function to maximize the long-term reward of the main patrol mission.

2. The system according to claim 1, characterized in that, The low-level strategies include restricted area avoidance strategy, static obstacle avoidance strategy, dynamic obstacle avoidance strategy, efficiency-first static obstacle avoidance strategy, and smooth action strategy; the intrinsic reward of the safety factor for each low-level strategy is as follows: Restricted Area Avoidance Strategy: When a patrol robot enters a pre-set restricted area, it will receive a huge negative reward; Static obstacle avoidance strategy: The reward is inversely proportional to the distance between the patrol robot and the nearest static obstacle; the closer the distance, the greater the negative reward. Dynamic obstacle avoidance strategy: The reward is inversely proportional to the distance between the patrol robot and the nearest dynamic individual, and the penalty coefficient is much higher than that of static obstacle avoidance; Smooth motion strategy: Negative rewards are given for the patrol robot’s sharp turns, rapid acceleration and deceleration to encourage smooth and predictable movements; Efficiency-first static obstacle avoidance strategy: This strategy enables patrol robots to prioritize task efficiency while avoiding static obstacles, allowing them to patrol at higher speeds in open areas.

3. The system according to claim 1, characterized in that, The adaptive security meta-policy is trained using the policy gradient method, and the reward it receives is the extrinsic reward obtained from actually performing the task, i.e., the reward for the main patrol task. Adaptive safety meta-policies learn to discover and understand that following different low-level policies in different contexts is necessary to better accomplish the final task.

4. The system according to claim 1, characterized in that, The safety factor reward function is shown in the following formula: Where, λ danger The intensity of the penalty for danger is represented by a safety factor genetic algorithm that evolves; D thresh is the hazard activation threshold, s represents the current state information perceived by the patrol robot from the environment at a certain point in time, a represents the action performed by the patrol robot in state s, and s' represents the new state that the patrol robot transitions to after performing action a in state s; II(*) is the indicator function, which is 1 when condition * is true, and 0 otherwise; D(s,a) represents the potential hazard assessment function. The relevant features representing the state-action pairs of the patrol robot are defined as phi_i(s,a). In the safety factor genetic algorithm, the chromosome encoded by the gene is a real-number vector containing the weights (w_1,...,w_N) of each feature and the intensity λ of the danger penalty. danger N is the total number of features; Using each chromosome generated by the safety factor genetic algorithm as a set of parameters, perform the following steps: A1. Use this set of parameters to construct the safety factor reward function R. SF ; A2. Select a standard reinforcement learning algorithm to train the low-level policy or the adaptive safety meta-policy; A3. Use the total reward R_total = r_task + R SF The patrol robot is trained for M episodes for evaluation; the trained patrol robot is evaluated, and the fitness function F is calculated; the fitness function F aims to maximize task utility while minimizing actual constraint violations and prediction errors.

5. The system according to claim 4, characterized in that, phi_i(s,a) includes: phi_1 represents the reciprocal of the Euclidean distance between the patrol robot and the nearest static obstacle after the robot performs action 'a'; a larger value indicates a closer distance and a higher level of danger. phi_2 represents the reciprocal of the Euclidean distance between the patrol robot and the nearest dynamic individual after the robot performs action 'a'; a larger value indicates a closer distance and a higher level of danger. phi_3 indicates whether the patrol robot has entered a known restricted area after performing action a. If it has entered the restricted area, the value is 1; otherwise, it is 0. phi_4 represents the normalized magnitude of the patrol robot's current linear velocity; phi_5 represents the normalized magnitude of the patrol robot's current angular velocity; phi_6 represents the reciprocal of the distance between the patrol robot and the least recently visited patrol point.

6. The system according to claim 4, characterized in that, Define task utility J util Define the average task utility return for the patrol robot during the evaluation phase; define the actual constraint violation J. viol_actual Let J_danger_prediction_accuracy be the actual average constraint violation cost of the patrol robot during the evaluation phase; the accuracy of hazard prediction (J_danger_prediction_accuracy) is defined as the accuracy of evaluating the prediction of D(s,a), specifically by calculating the activation of D(s,a), i.e., D(s,a) exceeds D_0. thresh However, the false positive rate when no actual violation occurred, and the false negative rate when an actual violation occurred but D(s,a) was not activated; The fitness formula is: F = W util *J util -W viol *J viol_actual -W pre_error *(FalsePositiveRate+FalseNegativeRate) Wherein, fitness F is the final score used in the genetic algorithm to evaluate the quality of each group's safety factors, and is used to maximize the F value; FalsePositiveRate represents the false alarm rate, which is the proportion of the potential hazard assessment function D(s,a) being activated, i.e., D(s,a) exceeding the threshold, but no violation actually occurring; FalseNegativeRate represents the false negative rate, which is the proportion of violations actually occurring, but D(s,a) not being activated; weight W util W viol and W pre_error These are all hyperparameters that need to be adjusted through training to balance the three objectives of task completion efficiency, safety efficiency, and prediction accuracy.

7. The system according to claim 4, characterized in that, The genetic operators in the safety factor genetic algorithm use crossover with standard real number encoding.

8. A method for a patrol robot to operate using the hierarchical adaptive security reinforcement learning system described in any one of claims 1-7, characterized in that, The specific process is as follows: The patrol robot perceives current status information from the environment; The high-level adaptive safety meta-policy receives the current state information perceived by the patrol robot as input, and then selects the most suitable low-level policy from the underlying safety factor policy pool based on the current environmental state. The selected low-level policy outputs a navigation action based on its learned behavioral preferences. The patrol robot executes the navigation action output by the low-level policy, interacts with the environment, the environmental state changes, and an immediate reward r_task and a potential safety factor penalty are generated. The total reward signal R_total, obtained by summing the immediate reward r_task and the potential safety factor penalty, is R_total = R_task + R SF This is further used to train an adaptive security meta-policy, enabling it to learn when to choose which security preferences while maximizing rewards for major patrol missions in the long run.

Citation Information

Patent Citations

  • Fixed-wing unmanned aerial vehicle autonomous control cooperation strategy training method

    CN112034888A

  • Multi-robot safety navigation method and system based on hierarchical deep reinforcement learning

    CN116339331A

  • Intelligent motion control method based on SAC reinforcement learning algorithm

    CN118311878A

  • Database adaptive data flow acquisition optimization method and system based on reinforcement learning

    CN119719783A

  • Path planning method based on machine learning in easily degraded environment

    CN119845286A