Safety sensing method for working environment of AI-driven robot
Through multimodal sensors and deep learning algorithms, multi-dimensional perception information is generated and behavioral strategies are dynamically adjusted, solving the problem of robot perception and response delay in complex environments, and achieving more efficient and safe robot operations.
Patent Information
- Application Number
- CN202510324520.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to achieve robots' rapid and accurate perception and understanding of the environment in complex environments, especially in limited space and extreme environments, and there is a delay in processing real-time data and dynamic environment changes, making it difficult to meet the needs of rapid response.
Multimodal sensors are used to collect visual, lidar, force awareness and environmental parameter data in real time, use deep learning algorithms to extract feature and semantic understanding, generate multi-dimensional perception information, and predict and classify potential risks based on preset security strategy models, and dynamically adjust the robot's behavioral strategy. Through the continuous learning module online optimization of deep learning models and security policy models to adapt to environmental changes.
It realizes the robot's more comprehensive and accurate perception and robustness in complex environments, and can respond quickly to environmental changes, ensuring efficient and safe in dynamic environments.
Smart Images

Figure CN120190815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically relates to an AI-driven safety perception method for a robot operating environment. Background Art
[0002] With the continuous development of industrial automation and intelligence, robots are increasingly widely used in various operating environments. To ensure the safety and efficiency of robots during operation, a method capable of real-time perceiving and evaluating environmental safety is required. The AI-driven safety perception method for a robot operating environment has emerged, which uses artificial intelligence technology to process and analyze data collected by sensors, thereby realizing real-time monitoring and early warning of environmental safety.
[0003] A single sensor (such as a camera or lidar) is easily interfered with in a complex environment and cannot provide comprehensive environmental information. In an unstructured and dynamically changing environment, it is difficult for a robot to quickly and accurately perceive and understand the environment, especially in a limited space and extreme environment; there are still delays in processing real-time data and dynamic environmental changes in the prior art, making it difficult to meet the requirements of rapid response. The fusion of different types of sensor data needs to solve problems such as data synchronization and error correction, and the current technology still needs to be improved in this regard; robots still have deficiencies in understanding complex environmental semantics and human instructions, making it difficult to achieve true intelligent interaction. Summary of the Invention
[0004] To solve the above technical problems, an AI-driven safety perception method for a robot operating environment is provided. This technical solution solves the problem that a single sensor is easily interfered with in a complex environment and cannot provide comprehensive environmental information. In an unstructured and dynamically changing environment, it is difficult for a robot to quickly and accurately perceive and understand the environment, especially in a limited space and extreme environment; there are still delays in processing real-time data and dynamic environmental changes, making it difficult to meet the requirements of rapid response. The fusion of different types of sensor data needs to solve the problems of data synchronization and error correction.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: An AI-driven safety perception method for a robot operating environment, comprising: Real-time collecting visual, lidar, force sense, and environmental parameter data of the operating environment through a multi-modal sensor; Using a deep learning algorithm to extract features and understand semantics from the collected data, and generating multi-dimensional perception information of the environment; Based on a preset safety policy model, combining multi-dimensional perception information, predicting and classifying potential risks; Dynamically adjust the robot's behavior strategy according to the risk classification results to achieve active safety decision-making; Through the continuous learning module, use the newly collected data to online optimize the deep learning model and the safety policy model to adapt to environmental changes.
[0006] Preferably, through multi-modal sensors, visually, lidar, force sense, and environmental parameter data of the operating environment are collected in real time, specifically including: Obtain different modal data in real time from multiple sensors, including collecting images or videos through a camera to obtain the three-dimensional point cloud information of the environment, collecting contact force information through force and tactile sensors, and collecting environmental parameters through temperature and humidity sensors; Preprocess the collected multi-modal data, convert the data of different modalities to a unified scale, remove the noise and invalid information in the data, and ensure that the multi-modal data is within the same reference framework through time synchronization and spatial registration.
[0007] Preferably, use deep learning algorithms to extract features and semantic understanding from the collected data to generate multi-dimensional perception information of the environment, specifically including: Based on the data collected by multi-modal sensors, use convolutional neural networks, point cloud processing networks, and time series analysis methods to extract features from image data, three-dimensional point cloud data, and force sense data respectively to obtain corresponding feature vectors; Concatenate and weighted sum the above-extracted feature vectors to obtain a fused feature; Based on the fused feature, use object detection algorithms to identify objects in the environment, use semantic segmentation networks to perform pixel-level and point-level segmentation on image and point cloud data, use deep learning models to predict the motion trajectories of objects, and use generative models to construct a three-dimensional model of the environment; Generate multi-dimensional perception information of the environment, including the objects in the environment and their position information, the motion trajectories and speeds of the objects, the three-dimensional structure and semantic segmentation results of the scene, potential collision risks, and dangerous areas.
[0008] Preferably, concatenate and weighted sum the above-extracted feature vectors to obtain a fused feature, specifically including: ; In the formula, , , are the feature vectors of image data, three-dimensional point cloud data, and force sense data respectively; α, β, γ are weight coefficients used to adjust the contributions of different modal features.
[0009] Preferably, based on a preset safety policy model, combined with multi-dimensional perception information, predict and classify potential risks, specifically including: Use deep learning algorithms to analyze the fused features and predict potential risks; Classify potential risks into different levels and classify them in combination with a preset security policy model; Evaluate the severity and priority of risks through methods such as risk matrices and risk maps.
[0010] Preferably, according to the risk classification results, dynamically adjust the robot's behavior strategy to achieve active safety decision-making, which specifically includes: According to the risk classification results, select corresponding behavior strategies from a preset security policy library; When actual risks are detected, combine global path planning and local path optimization to re-plan the path in real time to avoid dangerous areas; According to the risk type and level, execute corresponding obstacle avoidance strategies, and dynamically adjust the strategy parameters according to the execution results and environmental changes to optimize the behavior decision-making; During the execution process, continuously monitor environmental changes, update the risk assessment in real time, and dynamically adjust the behavior strategy according to the new risk information.
[0011] Preferably, when actual risks are detected, combine global path planning and local path optimization to re-plan the path in real time to avoid dangerous areas, which specifically includes: Global path planning, the A* algorithm includes: ; In the formula, f(n) is the total cost estimate of node n; g(n) is the actual cost from the starting point to node n; h(n) is the heuristic function, estimating the cost from node n to the target point; Local path optimization, the dynamic window method includes: ; In the formula, is the velocity vector, including linear velocity and angular velocity; V is the feasible velocity space; cost(v) is the cost function of the velocity vector, including target proximity and path safety.
[0012] Preferably, through a continuous learning module, use newly collected data to online optimize the deep learning model and the security policy model to adapt to environmental changes, which specifically includes: Use online learning and incremental learning methods to dynamically update the deep learning model; According to the new risk assessment results, dynamically adjust the threshold of the security policy model; By defining a reward function, enable the model to learn behaviors that meet safety standards, and perform closed-loop reinforcement learning fine-tuning in a simulated environment to obtain an optimized model; Perform real-time evaluation on the optimized model to verify its performance in the new environment. According to the evaluation results, further adjust the model parameters and optimization strategies; Adopt a closed-loop feedback mechanism, use the results of each optimization as new learning materials, and continuously improve the model.
[0013] Preferably, by defining a reward function, ensure that the model learns behaviors that meet safety standards, and perform closed-loop reinforcement learning fine-tuning in a simulation environment to obtain the optimized model, specifically including: In the reward function, the cumulative discounted reward is used to evaluate the long-term value of the behavior sequence: ; In the formula, is the cumulative discounted reward starting from time step t; γ is the discount factor, and the range is 0 ≤ γ < 1; is the reward obtained at time step t + k + 1.
[0014] Preferably, by defining a reward function, ensure that the model learns behaviors that meet safety standards, and perform closed-loop reinforcement learning fine-tuning in a simulation environment to obtain the optimized model, specifically including: In closed-loop reinforcement learning, the use of GRPO optimization objectives includes: ; In the formula, G is the optimization objective, is the probability of taking action under the current policy in state s; is the probability distribution under the old policy; is the advantage value of action ; λ is a hyperparameter that controls the weight of the KL divergence.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes to collect environmental data in real time through multi-modal sensors, and use deep learning algorithms for feature extraction and semantic understanding, which can generate more comprehensive and accurate multi-dimensional perception information, overcome the limitations of single sensors, and improve the perception ability and robustness of robots in complex environments; through the continuous learning module, use the newly collected data to online optimize the deep learning model and the safety policy model to adapt to environmental changes, only improve the adaptability and generalization ability of the model, and continuously improve the model performance through the closed-loop feedback mechanism; real-time data processing and feedback can quickly respond to environmental changes, the robot continuously monitors the environment during the task execution, and dynamically adjusts the behavior strategy according to the new risk information to ensure high efficiency and safety in the dynamic environment at all times; combines a variety of advanced technologies and algorithms, and can effectively cope with complex and changeable environments, such as mines, industrial workshops, medical scenarios, etc. Description of the Drawings
[0016] Figure 1 It is a flowchart of a method for safety perception in an AI - driven robot working environment; Figure 2 It is a flowchart of a method for multi - modal sensor data acquisition; Figure 3 It is a flowchart of a method for generating multi - dimensional perception information of the environment; Figure 4 It is a flowchart of a method for predicting and classifying potential risks; Figure 5 It is a flowchart of a method for dynamically adjusting the behavior strategy of the robot; Figure 6 It is a flowchart of a method for online optimization of the deep learning model and the safety policy model. Detailed implementation manners
[0017] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and other obvious variations can be thought of by those skilled in the art.
[0018] Refer to Figure 1 As shown, an AI - driven method for safety perception in a robot working environment includes: Real - time collect visual, lidar, force sense, and environmental parameter data of the working environment through multi - modal sensors; Use deep learning algorithms to extract features and perform semantic understanding on the collected data to generate multi - dimensional perception information of the environment; Based on a preset safety policy model, combine multi - dimensional perception information to predict and classify potential risks; According to the risk classification result, dynamically adjust the behavior strategy of the robot to achieve active safety decision - making; Through a continuous learning module, use the newly collected data to perform online optimization on the deep learning model and the safety policy model to adapt to environmental changes.
[0019] In this solution, by combining visual sensors (such as cameras), lidar, force sensors, and environmental parameter sensors (such as temperature and humidity sensors), the robot can obtain more comprehensive and accurate environmental information. For example, visual sensors are used to identify the shape and color of objects, lidar is used to measure distances and construct 3D maps, and force sensors are used to sense contact forces. The collaborative work of these sensors significantly improves the adaptability of the robot in complex environments.
[0020] In industrial automation, multi-modal sensors enable robots to perceive the position and state of objects on the production line in real time, thus achieving high-precision assembly and handling tasks; in the field of service robots, multi-modal perception technology enables robots to navigate autonomously in complex environments (such as hospitals and shopping malls) while avoiding collisions.
[0021] Refer to Figure 2 As shown, through multi-modal sensors, visual, lidar, force, and environmental parameter data of the operating environment are collected in real time, specifically including: Obtain different modal data in real time from multiple sensors, including collecting images or videos through cameras to obtain three-dimensional point cloud information of the environment, collecting contact force information through force and tactile sensors, and collecting environmental parameters through temperature and humidity sensors; Preprocess the collected multi-modal data, convert data of different modalities to a unified scale, remove noise and invalid information in the data, and ensure that the multi-modal data is within the same reference framework through time synchronization and spatial registration.
[0022] Refer to Figure 3 As shown, according to an AI-driven robot operating environment safety perception method described in claim 2, it is characterized in that the use of deep learning algorithms to extract features and understand semantics from the collected data to generate multi-dimensional perception information of the environment specifically includes: Based on the data collected by multi-modal sensors, use convolutional neural networks, point cloud processing networks, and time series analysis methods to extract features from image data, three-dimensional point cloud data, and force data respectively to obtain corresponding feature vectors; Concatenate and weighted sum the above-extracted feature vectors to obtain a fused feature; Based on the fused feature, use object detection algorithms to identify objects in the environment, use semantic segmentation networks to perform pixel-level and point-level segmentation on image and point cloud data, use deep learning models to predict the movement trajectories of objects, and use generative models to construct a three-dimensional model of the environment; Generate multi-dimensional perception information of the environment, including objects in the environment and their position information, the movement trajectories and speeds of objects, the three-dimensional structure and semantic segmentation results of the scene, potential collision risks, and dangerous areas.
[0023] The above-mentioned concatenation and weighted summation of the extracted feature vectors to obtain a fused feature specifically includes: ; In the formula, , , are the feature vectors of image data, three-dimensional point cloud data, and force data respectively; α, β, γ are weight coefficients used to adjust the contributions of different modal features.
[0024] It should be noted that the convolutional neural network (CNN) is used for image feature extraction and target recognition, and the recurrent neural network (RNN) and its variants (such as LSTM) are used to process time series data and predict the motion trajectory of objects. In the logistics scenario, deep learning algorithms enable robots to identify the location and category of goods in real time and predict their motion trajectories, thereby optimizing path planning and avoiding collisions. In the medical field, deep learning algorithms combined with multi-modal perception technology enable robots to identify the actions and needs of patients and provide precise auxiliary services.
[0025] Referring to Figure 4 As shown, the prediction and classification of potential risks based on a preset safety policy model and combined with multi-dimensional perception information specifically include: Using deep learning algorithms to analyze the fused features and predict potential risks; Dividing potential risks into different levels and classifying them in combination with a preset safety policy model; Evaluating the severity and priority of risks through methods such as risk matrices and risk maps.
[0026] Based on a preset safety policy model, the robot can predict and classify potential risks by combining multi-dimensional perception information, and evaluate the severity and priority of risks through risk matrices and maps. The robot can take measures in advance to avoid accidents. For example, by analyzing the motion trajectory and speed of an object, the robot can predict the possibility of a collision and adjust its behavior strategy according to the risk level. In high-risk environments such as mines, the robot monitors the collapse risk in real time through multi-modal perception and deep learning algorithms and evacuates the dangerous area in advance. In industrial scenarios, the robot can identify early signs of equipment failures and take maintenance measures in advance to reduce downtime.
[0027] Referring to Figure 5 As shown, the dynamic adjustment of the robot's behavior strategy according to the risk classification results to achieve active safety decision-making specifically includes: Selecting the corresponding behavior strategy from a preset safety policy library according to the risk classification results; When an actual risk is detected, combining global path planning and local path optimization to re-plan the path in real time to avoid dangerous areas; According to the risk type and level, execute the corresponding obstacle avoidance strategy, and dynamically adjust the strategy parameters according to the execution results and environmental changes to optimize the behavior decision-making; During the execution process, continuously monitor environmental changes, update the risk assessment in real time, and dynamically adjust the behavior strategy according to the new risk information.
[0028] When an actual risk is detected, combining global path planning and local path optimization to re-plan the path in real time to avoid dangerous areas specifically includes: Global path planning, the A* algorithm includes: ; In the formula, f(n) is the total cost estimate of node n; g(n) is the actual cost from the starting point to node n; h(n) is the heuristic function, estimating the cost from node n to the target point; Local path optimization, the dynamic window method includes: ; In the formula, is the velocity vector, including linear velocity and angular velocity; V is the feasible velocity space; cost(v) is the cost function of the velocity vector, including target proximity and path safety.
[0029] It should be noted that in high-risk environments such as mines, the robot monitors the collapse risk in real time through multi-modal perception and deep learning algorithms and evacuates the dangerous area in advance; in industrial scenarios, the robot can identify early signs of equipment failure and take maintenance measures in advance to reduce downtime; In the field of autonomous driving, the robot can drive safely in complex traffic environments by dynamically adjusting its behavior strategy; in logistics distribution, the robot significantly improves the distribution efficiency and safety through real-time path planning and obstacle avoidance strategies.
[0030] Refer to Figure 6 As shown, the online optimization of the deep learning model and the safety policy model by using the newly collected data through the continuous learning module to adapt to environmental changes specifically includes: Using online learning and incremental learning methods to dynamically update the deep learning model; According to the new risk assessment results, dynamically adjust the threshold of the safety policy model; By defining a reward function, enabling the model to learn behaviors that meet safety standards, and performing closed-loop reinforcement learning fine-tuning in a simulated environment to obtain an optimized model; Perform real-time evaluation on the optimized model, verify its performance in the new environment, and further adjust the model parameters and optimization strategies according to the evaluation results; Adopt a closed-loop feedback mechanism, use the results of each optimization as new learning materials, and continuously improve the model.
[0031] The specific process of ensuring that the model learns behaviors that meet safety standards by defining a reward function and performing closed-loop reinforcement learning fine-tuning in a simulated environment to obtain an optimized model includes: In the reward function, the cumulative discounted reward is used to evaluate the long-term value of the behavior sequence: ; wherein, is the cumulative discounted reward starting from time step t; γ is the discount factor with a range of 0 ≤ γ < 1; is the reward obtained at time step t + k + 1.
[0032] By defining the reward function to ensure that the model learns behaviors that meet safety standards, and performing closed-loop reinforcement learning fine-tuning in a simulation environment to obtain an optimized model specifically includes: In closed-loop reinforcement learning, the use of GRPO to optimize the objective includes: ; wherein, G is the optimization objective, is the probability of taking action under the current policy in state s; is the probability distribution under the old policy; is the advantage value of action ; λ is a hyperparameter that controls the weight of the KL divergence.
[0033] The continuous learning module of this solution enables the robot to online optimize the deep learning model and the safety policy model using newly collected data to adapt to environmental changes. Through reinforcement learning and a closed-loop feedback mechanism, the robot can be fine-tuned in a simulation environment and continuously optimize its behavior strategy in practical applications.
[0034] The usage process of the present invention is as follows: Step 1: Real-time obtain different modalities of data from multiple sensors, including collecting images or videos through a camera to obtain three-dimensional point cloud information of the environment, collecting contact force information through force and tactile sensors, and collecting environmental parameters through temperature and humidity sensors; Step 2: Preprocess the collected multi-modal data, convert data of different modalities to a unified scale, remove noise and invalid information in the data, and ensure that the multi-modal data is within the same reference framework through time synchronization and spatial registration; Step 3: Based on the data collected by the multi-modal sensors, respectively use a convolutional neural network, a point cloud processing network, and a time series analysis method to extract features from the image data, three-dimensional point cloud data, and force perception data to obtain corresponding feature vectors; Step 4: Concatenate and weighted sum the above-extracted feature vectors to obtain a fused feature; Step 5: Based on the fused feature, use an object detection algorithm to identify objects in the environment, use a semantic segmentation network to perform pixel-level and point-level segmentation on the image and point cloud data, use a deep learning model to predict the motion trajectory of the object, and use a generative model to construct a three-dimensional model of the environment; Step 6: Generate multi-dimensional perception information of the environment, including objects in the environment and their location information, the movement trajectories and speeds of the objects, the three-dimensional structure and semantic segmentation results of the scene, potential collision risks, and dangerous areas; Step 7: Use deep learning algorithms to analyze the fused features and predict potential risks; Step 8: Classify the potential risks into different levels and classify them in combination with a preset safety policy model; Step 9: Evaluate the severity and priority of the risks through methods such as risk matrices and risk maps; Step 10: Select corresponding behavior strategies from a preset safety policy library according to the risk classification results; Step 11: When actual risks are detected, combine global path planning and local path optimization to re-plan the path in real time to avoid dangerous areas; Step 13: According to the risk type and level, execute the corresponding obstacle avoidance strategy. Dynamically adjust the strategy parameters according to the execution results and environmental changes to optimize the behavior decision-making; Step 14: Continuously monitor environmental changes during the execution process, update the risk assessment in real time, and dynamically adjust the behavior strategy according to the new risk information; Step 15: Use online learning and incremental learning methods to dynamically update the deep learning model; Step 16: Dynamically adjust the threshold of the safety policy model according to the new risk assessment results; Step 17: By defining a reward function, enable the model to learn behaviors that meet safety standards, perform closed-loop reinforcement learning fine-tuning in a simulated environment, and obtain an optimized model; Step 18: Conduct real-time evaluation of the optimized model, verify its performance in the new environment, and further adjust the model parameters and optimize the strategy according to the evaluation results; Step 19: Adopt a closed-loop feedback mechanism, use the results of each optimization as new learning materials, and continuously improve the model.
[0035] In summary, the advantages of the present invention are as follows: By means of multi-modal sensors, environmental data is collected in real time, and deep learning algorithms are used for feature extraction and semantic understanding, enabling the generation of more comprehensive and accurate multi-dimensional perception information, overcoming the limitations of single sensors, and enhancing the perception ability and robustness of the robot in complex environments; Through the continuous learning module, the deep learning model and the safety policy model are optimized online using newly collected data to adapt to environmental changes, only improving the adaptability and generalization ability of the model, and continuously improving the model performance through a closed-loop feedback mechanism; Real-time data processing and feedback can quickly respond to environmental changes. The robot continuously monitors the environment during the task execution process and dynamically adjusts the behavior strategy according to new risk information to ensure high efficiency and safety in a dynamic environment at all times; Combining a variety of advanced technologies and algorithms can effectively cope with complex and changeable environments, such as mines, industrial workshops, medical scenarios, etc.
[0036] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-driven robot working environment safety perception method, characterized in that: include: Through multimodal sensors, real-time data on vision, lidar, force perception and environmental parameters of the working environment are collected; Use deep learning algorithms to extract features and understand semantics of collected data to generate multi-dimensional perception information of the environment; Based on the preset security policy model and combined with multi-dimensional perception information, potential risks are predicted and classified; According to the risk classification results, the robot's behavior strategy is dynamically adjusted to achieve proactive safety decision-making; Through the continuous learning module, the deep learning model and security policy model are optimized online using newly collected data to adapt to environmental changes.
2. According to claim 1, an AI-driven robot working environment safety perception method is characterized in that: The real-time collection of visual, laser radar, force perception and environmental parameter data of the working environment through multimodal sensors specifically includes: Acquire data of different modalities from multiple sensors in real time, including collecting images or videos through cameras, obtaining 3D point cloud information of the environment, collecting contact force information through force and tactile sensors, and collecting environmental parameters through temperature and humidity sensors; The collected multimodal data are preprocessed to convert data of different modes to a unified scale, remove noise and invalid information in the data, and ensure that the multimodal data are in the same reference frame through time synchronization and spatial alignment.
3. According to claim 2, an AI-driven robot working environment safety perception method is characterized in that: The use of deep learning algorithms to extract features and understand semantics of the collected data and generate multi-dimensional perception information of the environment specifically includes: Based on the data collected by multimodal sensors, convolutional neural networks, point cloud processing networks and time series analysis methods are used to extract features from image data, three-dimensional point cloud data and force data to obtain corresponding feature vectors. The feature vectors extracted above are concatenated and weighted summed to obtain fusion features; Based on the fusion features, the target detection algorithm is used to identify objects in the environment, the semantic segmentation network is used to perform pixel-level and point-level segmentation on images and point cloud data, the deep learning model is used to predict the movement trajectory of objects, and the generative model is used to build a three-dimensional model of the environment; Generate multi-dimensional perception information of the environment, including objects in the environment and their location information, object movement trajectory and speed, three-dimensional structure and semantic segmentation results of the scene, potential collision risks and dangerous areas.
4. According to claim 3, an AI-driven robot working environment safety perception method is characterized in that: The step of concatenating and weighted summing the feature vectors extracted above to obtain fusion features specifically includes: ; In the formula, , , are the feature vectors of image data, 3D point cloud data and force data respectively; α, β, γ are weight coefficients used to adjust the contribution of different modal features.
5. According to claim 4, an AI-driven robot working environment safety perception method is characterized in that: The prediction and classification of potential risks based on the preset security policy model and combined with multi-dimensional perception information specifically include: Use deep learning algorithms to analyze the fused features and predict potential risks; Classify potential risks into different levels and classify them according to the preset security policy model; Assess the severity and priority of risks through risk matrix and risk map.
6. The AI-driven robot working environment safety perception method according to claim 5 is characterized in that: The method of dynamically adjusting the robot's behavior strategy according to the risk classification results to achieve active safety decision-making specifically includes: According to the risk classification results, select the corresponding behavior strategy from the preset security strategy library; When actual risks are detected, the path is replanned in real time to avoid dangerous areas by combining global path planning and local path optimization; According to the risk type and level, the corresponding obstacle avoidance strategy is executed. According to the execution results and environmental changes, the strategy parameters are dynamically adjusted to optimize the behavior decision; During the execution process, environmental changes are continuously monitored, risk assessments are updated in real time, and behavioral strategies are dynamically adjusted based on new risk information.
7. The AI-driven robot working environment safety perception method according to claim 6 is characterized in that: When an actual risk is detected, the path is replanned in real time to avoid the dangerous area by combining global path planning and local path optimization. Specifically, the following steps are performed: Global path planning, A* algorithm includes: ; Where f(n) is the total cost estimate of node n; g(n) is the actual cost from the starting point to node n; h(n) is the heuristic function that estimates the cost from node n to the target point; Local path optimization, dynamic window method includes: ; In the formula, is the velocity vector, including linear velocity and angular velocity; V is the feasible velocity space; cost(v) is the cost function of the velocity vector, including target proximity and path safety.
8. The AI-driven robot working environment safety perception method according to claim 7, characterized in that: The continuous learning module uses newly collected data to perform online optimization on the deep learning model and security policy model to adapt to environmental changes, specifically including: Use online learning and incremental learning methods to dynamically update deep learning models; Dynamically adjust the threshold of the security policy model based on new risk assessment results; By defining a reward function, the model learns behaviors that meet safety standards, and performs closed-loop reinforcement learning fine-tuning in a simulation environment to obtain an optimized model; Conduct real-time evaluation of the optimization model to verify its performance in the new environment, and further adjust the model parameters and optimization strategy based on the evaluation results; A closed-loop feedback mechanism is adopted to use the results of each optimization as new learning material to continuously improve the model.
9. The AI-driven robot working environment safety perception method according to claim 8, characterized in that: By defining the reward function, ensuring that the model learns behaviors that meet safety standards, and performing closed-loop reinforcement learning fine-tuning in a simulation environment, the optimized model is obtained, specifically including: In the reward function, the cumulative discounted reward is used to evaluate the long-term value of the action sequence: ; In the formula, is the cumulative discounted reward starting from time step t; γ is the discount factor, ranging from 0≤γ<1; is the reward obtained at time step t+k+1.
10. The AI-driven robot working environment safety perception method according to claim 8, characterized in that: By defining the reward function, ensuring that the model learns behaviors that meet safety standards, and performing closed-loop reinforcement learning fine-tuning in a simulation environment, the optimized model is obtained, specifically including: In closed-loop reinforcement learning, the optimization objectives using GRPO include: ; In the formula, G is the optimization target, To take action in state s under the current policy probability; is the probability distribution under the old strategy; For Action The advantage value of ; λ is a hyperparameter that controls the weight of KL divergence.
Citation Information
Cited By
Robot dynamic risk assessment and decision-making system and method based on multi-modal perception
CN120680531A
Robot walking control method, device and equipment and medium
CN120909328A
Automatic closing robot control system for gas cylinder leakage in dangerous chemical accident site
CN121245845A
Humanoid robot risk early warning method and related equipment
CN121327585A
Humanoid robot risk warning method and related device
CN121327585B