Action strategy security enhancement method and device, equipment and medium

By constructing a safety constraint space and generating safe action strategies using multimodal perception features, the problem of lack of safety constraints in robot action strategies in existing technologies is solved, thus ensuring the safety and reliability of robots in complex environments.

CN120848207APending Publication Date: 2025-10-28PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511060285.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The existing control policy module lacks dynamic perception and effective constraints on safety risks in the operating environment, which makes the robot prone to safety problems such as collisions and path intrusion when performing actions. Especially in open and complex environments, it is unable to effectively integrate multi-source environmental information and safety constraints.

Method used

Construct a safety constraint space that includes physical safety constraints, environmental safety constraints, and preset safety conditions; acquire multimodal perception data to generate fused safety perception features; generate initial action strategies based on task objectives and project them onto the safety constraint space; monitor status data and trigger intervention measures during execution; and collect and update the strategy generation module.

Benefits of technology

It enables robots to dynamically ensure safety in complex environments. By fusing multimodal perception with task objectives, it generates safety action strategies, monitors and adjusts actions in real time to avoid safety risks, and improves the safety and reliability of robots in various business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848207A_ABST
    Figure CN120848207A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as agent autonomous decision making, financial science and technology and medical health, and discloses an action strategy security enhancement method, device and equipment and a medium. And generating an initial action strategy through a strategy generation module based on the task target and the fused security perception feature, projecting the initial action strategy to a security constraint space to obtain a security action strategy, monitoring state data of the controlled object in an execution process, and executing an intervention measure when the state data triggers a monitoring threshold. And collecting execution process data and updating a strategy generation module based on the execution process data. According to the method, the initial action strategy is generated through fusion of the multi-modal sensing information and the task target, and the safety of task execution of the robot in a complex environment is improved by combining the safety constraint space projection, the state monitoring and the data feedback updating strategy generation module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for enhancing the safety of action strategies. Background Technology

[0002] In existing research on robot control technology, control policy modules are typically task-oriented, focusing on efficiently completing designated tasks. However, these modules often lack dynamic perception and effective constraints on safety risks in the operating environment, easily leading to robots neglecting human safety and environmental risks when performing actions. For example, in physical interaction tasks, control policy modules struggle to adjust action decisions in real time according to environmental changes, easily resulting in safety issues such as accidental collisions and path intrusions. This traditional approach, prioritizing task completion efficiency, leads to significant shortcomings in robot behavior regarding safety, especially in open and complex environments, lacking the ability to integrate and process multi-source environmental information and safety constraints.

[0003] In the healthcare sector, with the increasing application of medical robots in patient care and surgical assistance, existing control policy modules still primarily focus on surgical path planning or task objective achievement, lacking sufficient integration of multimodal perception data such as patient status, spatial distance, and environmental risks for safety considerations. As a result, when robots assist in moving patients or performing medical procedures, the lack of real-time perception of changes in patient posture and the clinical environment may lead to collisions or erroneous actions, posing a potential threat to patient safety.

[0004] In the fintech sector, intelligent robots are used for tasks such as guidance and item delivery in financial service venues. However, existing control policy modules fail to effectively integrate the dynamic flow of people and safety constraints in the service environment. This can lead to situations where robots fail to proactively avoid customers during peak hours or intrude into customers' safe distances, impacting user experience and posing security risks. Especially in financial service outlets with dense crowds and complex dynamic environments, traditional modules cannot flexibly adjust their action strategies based on environmental perception information, lacking dynamic safety protection capabilities.

[0005] In summary, existing control policy modules generally suffer from a disconnect between task orientation and safety assurance. They cannot fully utilize multimodal perception data to comprehensively constrain physical attributes, environmental states, and preset safety conditions in the environment, making it difficult to effectively ensure the safety and reliability of robot action strategies in various business scenarios. Summary of the Invention

[0006] The main objective of this invention is to provide a method, apparatus, device, and storage medium for enhancing the safety of action strategies, aiming to solve the technical problem that traditional action strategy methods lack a full-process safety constraint mechanism, leading to safety risks such as collisions and misoperations during execution.

[0007] To achieve the above objectives, the present invention provides a method for enhancing the security of action strategies, comprising:

[0008] Construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions;

[0009] Acquire multimodal perception data and generate fused security perception features based on the multimodal perception data;

[0010] Based on the task objective and the fused security awareness features, an initial action policy is generated through the policy generation module.

[0011] The initial action strategy is projected onto the safety constraint space to obtain the safety action strategy;

[0012] During the execution of the safety action strategy, the status data of the controlled object is monitored, and when the status data triggers a preset monitoring threshold, intervention measures are executed.

[0013] Collect execution process data of the security action policy, and update the policy generation module based on the execution process data.

[0014] Furthermore, to achieve the above objectives, the present invention provides a motion strategy safety enhancement device, comprising:

[0015] The safety constraint modeling module is used to construct a safety constraint space that includes physical safety constraints based on the physical properties of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions.

[0016] A multimodal perception fusion module is used to acquire multimodal perception data and generate fused security perception features based on the multimodal perception data;

[0017] The strategy generation module is used to generate an initial action strategy based on the task objective and the fused security awareness features.

[0018] A security policy projection module is used to project the initial action policy onto the security constraint space to obtain a security action policy.

[0019] The safety monitoring and intervention module is used to monitor the status data of the controlled object during the execution of the safety action strategy, and to execute intervention measures when the status data triggers a preset monitoring threshold.

[0020] The policy adaptive optimization module is used to collect execution process data of the security action policy and update the policy generation module based on the execution process data.

[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an action policy security enhancement program stored in the memory and executable on the processor, wherein when the action policy security enhancement program is executed by the processor, it implements the steps of the action policy security enhancement method as described above.

[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an action policy security enhancement program, which, when executed by a processor, implements the steps of the action policy security enhancement method as described above.

[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as intelligent agent autonomous decision-making, fintech, and healthcare. It discloses a method, apparatus, device, and medium for enhancing the safety of action strategies, including: constructing a safety constraint space; acquiring multimodal perception data to generate fused safety perception features; generating an initial action strategy based on the task objective and the fused safety perception features through a strategy generation module; projecting the initial action strategy onto the safety constraint space to obtain a safe action strategy; monitoring the state data of the controlled object during the execution of the safe action strategy and implementing intervention measures when the state data triggers a monitoring threshold; collecting execution process data of the safe action strategy and updating the strategy generation module based on the execution process data. This invention, by fusing multimodal perception information with the task objective to generate an initial action strategy and projecting the initial action strategy onto a safety constraint space containing physical constraints, environmental constraints, and preset conditions, and by combining state data monitoring and execution process data feedback to update the strategy generation module, can dynamically ensure the safety of robots performing tasks in complex environments. Attached Figure Description

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0025] Figure 1 This is a schematic diagram of an application environment for the action strategy security enhancement method in one embodiment of the present invention;

[0026] Figure 2 This is a flowchart illustrating an embodiment of the motion strategy security enhancement method of the present invention;

[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the motion strategy safety enhancement device of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0029] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0030] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0031] The motion strategy security enhancement method provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can construct a safety constraint space through the user terminal, acquire multimodal perception data to generate fused safety perception features, generate an initial action policy based on the task objective and the fused safety perception features through a policy generation module, project the initial action policy onto the safety constraint space to obtain a safe action policy, monitor the state data of the controlled object during the execution of the safe action policy, and execute intervention measures when the state data triggers a monitoring threshold. The execution process data of the safe action policy is collected and the policy generation module is updated based on the execution process data. This invention dynamically ensures the safety of the robot performing tasks in complex environments by fusing multimodal perception information with the task objective to generate an initial action policy and projecting the initial action policy onto a safety constraint space containing physical constraints, environmental constraints, and preset conditions. Simultaneously, it combines state data monitoring and execution process data feedback to update the policy generation module. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the motion strategy security enhancement method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0033] like Figure 2 As shown, the motion strategy security enhancement method proposed in this invention includes the following steps:

[0034] S10, Construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions;

[0035] In this embodiment, in constructing a safety constraint space that includes physical safety constraints based on the physical properties of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions, the first step is to collect the mechanical structure parameters and motion characteristic parameters of the controlled object. Mechanical structure parameters include, but are not limited to, joint dimensions, link lengths, and the positional relationships of connecting nodes. These parameters can be obtained by reading mechanical design data using digital modeling software or by measuring them through laser scanning. Motion characteristic parameters include maximum joint rotation angle, maximum linear velocity, and acceleration threshold, which can be determined through experimental measurement or analysis of historical motion data. A mapping relationship is established between the mechanical structure parameters and the motion characteristic parameters to limit the robot's range of motion and physical load-bearing capacity, thereby forming physical safety constraints such as range of motion limits, load capacity limits, and movement speed limits. This part of the implementation requires combining multiple parameters into a unified coordinate system to ensure the comparability of data from different sources.

[0036] Meanwhile, environmental perception data is acquired based on multi-source sensor input, including environmental state information obtained from devices such as LiDAR, depth cameras, ultrasonic sensors, and inertial navigation units. This environmental perception data undergoes data fusion to extract spatial distribution information, forming an environmental feature matrix, specifically including obstacle distribution information and hazardous area location information. During processing, the environmental feature matrix is ​​spatially aligned through coordinate transformation and map modeling to accurately describe the environmental constraints associated with the controlled object in the external space, thereby generating environmental safety constraints.

[0037] Based on the acquired physical and environmental safety constraints, a safety condition library is accessed to obtain preset safety conditions. This library is a manually defined and maintained set of rules, including safety distance requirements, prohibited areas, and restrictions on dangerous behaviors. The preset safety conditions are uniformly encoded with the aforementioned physical and environmental safety constraints to form a mathematical constraint set. This mathematical constraint set describes constraint boundaries and conditions through inequalities and logical expressions, allowing it to be directly used in subsequent calculations. The entire safety constraint space, by jointly defining physical, environmental, and preset safety constraints, possesses the ability to comprehensively cover both internal and external environmental limitations of the controlled object, and provides a real-time, calculable constraint input interface.

[0038] Mechanical structural parameters and motion characteristic parameters can be collected using mechanical design software combined with finite element analysis tools, suitable for controlled objects with rigid structures. High-precision 3D scanning equipment and dynamic sensor networks can also be used to collect mechanical structural and motion characteristic parameters of flexible mechanisms and irregular structures, adapting to application scenarios such as flexible service robots. In terms of environmental perception data processing, dense point cloud modeling methods based on LiDAR can be used, suitable for real-time modeling needs in outdoor dynamic environments. Alternatively, SLAM mapping using depth camera and IMU data fusion can be used, adapting to indoor multi-obstacle environments. Regarding the configuration of preset safety conditions, parameterized configuration can be achieved through a rule engine to support dynamic adjustment of high-precision safety distance requirements in medical and health environments. A strategy template system can also be used to adapt to the safety prohibition requirements of financial service robots in multi-user interactive environments within business halls. The unified encoding of the safety constraint space can be represented in matrix form to combine multi-dimensional spatial constraints, enabling matrix operations with the multi-dimensional parameters of the action strategy.

[0039] Example: In the field of autonomous decision-making for intelligent agents, robots can collect data on the limits of the robotic arm joints and the gripping load capacity, and combine this with dynamic environmental perception to construct an adaptive safety constraint space, ensuring that autonomous path planning does not exceed structural limits and avoids environmental obstacles.

[0040] In the field of healthcare, surgical robots can form precise physical safety constraints based on the geometric features of surgical instruments and the physiological parameters of patients. At the same time, they can construct environmental safety constraints through spatial perception data of the surgical area, and define preset safety conditions by combining the patient's sensitive area protection rules, so as to ensure that the surgical operation is precise and safe.

[0041] In the fintech business, branch service robots can construct a safety constraint space by combining their own motion parameters with the perceived layout of the branch, and generate dynamic safety rules by analyzing customer behavior data, thereby avoiding potential physical contact risks to customers and improving the safety and reliability of services.

[0042] This embodiment forms physical safety constraints by collecting mechanical structure and motion characteristic parameters, generates environmental safety constraints by fusing environmental perception data, and combines them with preset safety conditions to form a unified set of mathematical constraints. This enables comprehensive coverage and computational support of physical constraints, environmental restrictions, and behavioral rules during the action decision-making process, thereby ensuring that the robot has multi-dimensional safety constraints when performing action decisions.

[0043] S20, acquire multimodal perception data, and generate fused security perception features based on the multimodal perception data;

[0044] In this embodiment, the process of acquiring multimodal perception data and generating fused security perception features based on this data first requires collecting multimodal perception data, covering multiple input sources such as vision, audio, spatial distance, and state information. Visual data mainly comes from cameras, including image frame sequences. This data is processed through convolutional neural networks to extract visual feature information including edges, textures, object shapes, and color distribution. Audio data comes from microphone arrays; after acquisition, spectral features are obtained through short-time Fourier transform, and then the frequency distribution and amplitude patterns are analyzed by a voiceprint feature extraction module. Spatial distance data comes from LiDAR, depth cameras, or ultrasonic sensors, requiring point cloud reconstruction to form spatial structure information containing spatial positional relationships and local density features. State information data includes acceleration and angular velocity data from the inertial measurement unit, as well as velocity and position sensor data from within the system. This data needs to be normalized and time-series synchronized to unify the sampling timing of data from different sources.

[0045] In generating fused security-aware features, it is necessary to unify and align the dimensions of visual features, audio features, spatial structure information, and state information data in the feature space. A multi-head self-attention mechanism is used to analyze the correlation between different modalities. Specifically, by assigning weight coefficients to each pair of cross-modal features and calculating their interdependence, the final fused features can enhance the overall expressive ability of multimodal information related to security. The fused security-aware features contain rich cross-modal contextual information, reflecting a comprehensive expression of objects, sounds, space, and motion states in the current environment, while retaining detailed features closely related to security risk assessment in each modality. This process ensures the temporal and spatial alignment consistency between different modal data and forms a highly comprehensive expression in the deep feature space.

[0046] Visual features can be extracted by acquiring RGB image sequences through a visual camera and combining edge detection and convolutional feature extraction methods for shape and texture recognition of static objects. They can also be combined with optical flow calculations between video frames to capture the trajectories of moving objects for real-time perception in dynamic environments. In audio data processing, single-channel microphones can be used to acquire audio features through Mel-spectrum mapping, suitable for voiceprint recognition in quiet environments. Alternatively, multi-microphone arrays can be used with beamforming algorithms to enhance sound source direction features for separation and analysis of multiple sound sources in complex environments. For spatial distance data processing, LiDAR can be used to construct dense point cloud maps, suitable for high-precision spatial reconstruction in large outdoor environments. Depth cameras combined with IMU depth compensation algorithms can be used for near-range obstacle detection in indoor environments. State information data can be acquired using low-latency inertial sensors to collect high-frequency acceleration and angular velocity data for real-time feedback of dynamic motion states. Alternatively, a combination of wheel speedometers and encoders can be used to obtain precise speed and position information for ground-based mobile robots. Multi-head self-attention mechanisms can use positional encoding to compensate for time-series information to preserve the temporal dynamics of motion states, or they can use embedding vectors to align cross-modal feature spaces to ensure consistency of fused features across feature dimensions.

[0047] Example: In the field of autonomous decision-making for intelligent agents, service robots can simultaneously collect environmental images and spatial point cloud data through cameras and LiDAR, and combine them with audio information and their own motion state data to generate cross-modal fusion perception features, providing accurate perception of surrounding dynamic obstacles and environmental changes for indoor autonomous navigation decisions.

[0048] In the healthcare business, surgical robots can acquire high-resolution images of the surgical field through endoscopic cameras, and combine them with real-time audio information of the surgical area and instrument movement data to form a fusion of safety perception features, which can be used to support a deep understanding of the tissue structure and operational risks in the surgical area.

[0049] In the fintech business, branch guidance robots can collect information on customer group distribution and communication status through vision and audio. Combined with their own movement status and spatial layout information, they can generate integrated security perception features, providing comprehensive perception support for autonomous path planning based on customer density, communication volume, and passable areas.

[0050] This embodiment acquires and deeply analyzes multimodal data from vision, audio, spatial distance, and state information. It aligns and associates environmental and object perception data from different sources in a unified space to form a fusion security perception feature containing rich context. This provides high-precision, multi-dimensional environmental understanding capabilities for action decision-making, enhances the perception foundation for subsequent action strategy generation, and strengthens the perception and expression of potential security risks in complex scenarios.

[0051] S30, Based on the task objective and the fused security awareness features, an initial action strategy is generated through the strategy generation module;

[0052] In this embodiment, during the generation of the initial action strategy through the strategy generation module based on the task objective and fused safety perception features, the task objective first needs to be parsed. The task objective can originate from high-level task instructions or user input, and includes the specific task type, execution parameters, and expected completion conditions. Parsing the task objective requires transforming the original task description into a format suitable for computational processing, such as converting the task content described in natural language into a task type identifier and parameter set. The task type defines the robot's behavior category, and the parameter set limits the specific scope and constraints of the behavior. The task priority and key operation points extracted from the task description will serve as important components of the task feature representation; these directly determine the action generation order and importance allocation of the subsequent decision-making module.

[0053] Task feature representation and the fusion of security-aware features require dimensional alignment and concatenation. Dimensional alignment requires matching the distribution characteristics and dimensions of both in the feature space, so that inputs from two different sources can be used as unified comprehensive input features, facilitating processing by the policy generation module. The resulting comprehensive input features not only contain task-related target semantic information but also incorporate contextual data from the environment that may affect security risks.

[0054] In the implementation of an agent's vision-language-action model, the policy generation module is often referred to as the control policy module. It is responsible for outputting robot action decisions based on task objectives and multimodal perception information. The policy generation module employs a reinforcement learning policy network. This network first uses a feature extraction layer to perform deep feature encoding on the comprehensive input features, extracting high-order representations that include task-environment relationships. The feature extraction process can use convolutional networks to process local feature patterns or recurrent neural networks to model time-series features, ensuring consistent temporal dependencies. The extracted features are then input into a fully connected layer to map the high-dimensional features to the action space, learning the probability distribution of different action types through a weight matrix. This probability distribution reflects the likelihood of the robot selecting various actions within the current task and environmental context.

[0055] The output layer further processes the motion space mapping, combining it with the prediction requirements of continuous-time motion parameters to generate a sequence of motion instructions. These time-series motion instructions not only indicate the motion type but also specify the execution time, duration, and sequence requirements for each motion. These motion instructions are encapsulated to form an initial motion strategy, serving as the direct input for robot control execution.

[0056] The task scheduling module can parse user-inputted text task descriptions and convert them into task type identifiers and accompanying parameters, making it suitable for structured data-driven industrial robot scheduling scenarios. The speech recognition and intent parsing module can also convert speech input into task commands, suitable for service robot applications. Task feature representation and fusion of safety-aware features can be achieved through dimension-aligned concatenation using embedding vectors, such as linear concatenation using feature embedding spaces of the same dimension, or by using an attention-weighted concatenation mechanism to enhance the expression of task importance weights. The feature extraction layer can employ multi-layer convolutional neural networks for feature compression and abstraction of image and point cloud types, or bidirectional long short-term memory networks for dependency modeling of time-series data. The fully connected layer can use multi-branch outputs to learn the selection probability distributions of multiple action types in parallel, suitable for multi-task concurrent execution scenarios. The output layer can further optimize the estimation accuracy of time-continuous action parameters using Bayesian time-series prediction algorithms, suitable for precise motion control of robots in highly dynamic environments.

[0057] Example description: In the field of autonomous decision-making for intelligent agents, by comprehensively analyzing the workshop scheduling target and the on-site environmental perception data, when generating the initial action strategy, the robot uses the task type and parameters in the task description as the main input factors, combined with the real-time collected visual and spatial distance information, to perform high-order feature encoding in the reinforcement learning policy network, ensuring that the output action command sequence meets the scheduling requirements while avoiding the risk of collision in the work area.

[0058] In the field of healthcare, rehabilitation robots analyze training type and intensity parameters for patients' rehabilitation tasks, and combine them with the patient's surrounding environment perception data, especially location information and dynamic posture data, to generate initial action strategies. This ensures that the force and angle of rehabilitation actions are within a safe range, dynamically adapts to the patient's current state, and reduces the probability of discomfort or danger during training.

[0059] In the field of fintech business, when faced with the goal of providing guidance services, the bank's welcoming robot analyzes the start and end points and priority parameters of the customer guidance task, and combines the environmental perception data of the branch with the real-time flow of people to generate an initial action strategy. This allows the robot's planned path and action details to dynamically avoid densely populated areas, while completing the customer guidance task in the shortest possible time, ensuring service safety and efficiency.

[0060] This embodiment combines task objectives with fused safety perception features, and uses a reinforcement learning policy network to generate action command sequences from feature encoding and spatial mapping. This enables action decision results to take into account both task requirements and environmental safety conditions, ensuring the adaptability, rationality, and safety of robot action strategies, reducing the probability of high-risk action commands, and improving the contextual awareness of action generation in complex task scenarios.

[0061] S40, Project the initial action strategy onto the safety constraint space to obtain the safety action strategy;

[0062] In this embodiment, the initial action strategy refers to the set of robot action decisions formed after comprehensive analysis of task objectives and environmental perception, which includes time-series action parameters.

[0063] The safety constraint space refers to a multi-dimensional set of constraints formed by combining the physical properties of the controlled object, external environmental perception, and predefined safety conditions, limiting the range of feasible solutions for robot motion parameters. The projection operation specifically compares the joint angle parameters, movement speed parameters, and load requirement parameters involved in the initial motion strategy with the motion range limits, movement speed limits, and load capacity limits contained in the safety constraint space. When any motion parameter is detected to exceed the constraint range, it is marked as a violation parameter and its boundary is corrected to converge to the boundary value of the constraint range or the tolerance range within the range. After correction, environmental safety constraints are introduced to further verify whether the motion correction result meets the feasibility conditions of the environmental space, such as obstacle avoidance and hazardous area elimination verification requirements, ensuring that the motion correction is not only effective in the parameter space but also executable in the environment.

[0064] Finally, the corrected motion parameters replace the corresponding non-compliant parameters in the original initial motion strategy, forming a safe motion sequence containing the corrected motion sequence, and then encapsulating it to generate a complete safe motion strategy. This process involves multi-dimensional matching and constraint adjustment of motion parameters with physical properties, environmental constraints, and task requirements, and has strict logical connections to ensure the safety consistency of robot actions in space, time, and dynamic behavior.

[0065] By defining an initial motion strategy data structure, including joint angle arrays, velocity arrays, and load parameter arrays, the system accesses in real-time the motion range limits, velocity thresholds, and load limits recorded in the safety constraint space. Joint angle parameters are compared using absolute values ​​and confined within physical motion limits; velocity parameters are truncated using a maximum velocity boundary truncation algorithm; and load parameters are matched using real-time maximum allowable load matching. If the environmental perception module outputs obstacle or danger zone markers, a spatial remapping algorithm is invoked to adjust the position and trajectory parameters involved in the motion sequence within the spatial geometric relationships. After adjustment, the adjusted parameters are reassembled into a new motion sequence and encapsulated as a safety motion strategy output. Flexible projection correction can be achieved through gradient projection algorithms, hard boundary correction can be configured through a rule engine, or dynamic adjustments can be made by combining soft constraint elastic tolerance strategies to adapt to the safety accuracy requirements of different tasks.

[0066] This embodiment uses the above processing to strictly limit the initial action strategy guided by the task objective within a multi-dimensional safety constraint space, ensuring that the robot's actions meet the comprehensive safety standards of physical safety, environmental safety, and task requirements in various working scenarios.

[0067] S50, during the execution of the safety action strategy, the status data of the controlled object is monitored, and when the status data triggers a preset monitoring threshold, intervention measures are executed;

[0068] In this embodiment, during the execution of the safety action strategy, the state data of the controlled object refers to the set of dynamic operating parameters collected during operation, including but not limited to position data, speed data, attitude data, and environmental distance data. These parameters reflect the real-time state of the controlled object in the time and space dimensions. Monitoring refers to the real-time sampling and analysis of the state data, comparing it with preset monitoring thresholds to determine whether there is an anomaly. The preset monitoring thresholds include a safety distance threshold, a movement speed threshold, and a load fluctuation threshold. These thresholds are derived from the modeling and experimental verification results of the safe operation characteristics of the controlled object and are used to define the boundaries of normal behavior. When any dimension parameter of the state data reaches or exceeds the corresponding threshold, the intervention decision process is immediately triggered. The intervention measures are real-time control responses executed according to the current anomaly type, divided into two types: operation warning and emergency interruption. Operation warning corresponds to minor anomalies, reducing risk by adjusting action parameters, such as decelerating and smoothing the movement curve. Emergency interruption corresponds to severe anomalies, protecting the system and environment by immediately stopping movement, activating safety protocols, and issuing safety alarms. Every step in the monitoring and intervention process is subject to strict parameter consistency constraints to ensure data closure and action consistency throughout the collection, comparison, judgment, and execution chain, thereby improving control reliability.

[0069] By configuring a multi-channel real-time data acquisition module, spatial position information, linear and angular velocities, attitude angles, distance sensor outputs, and load sensor outputs of the controlled object are collected at fixed time intervals. A comparison algorithm module sequentially compares environmental distance data with safe distance thresholds, speed data with motion speed thresholds, and load fluctuation amplitude with load fluctuation thresholds. A circular buffer queue can be used to cache short-term historical data windows to support real-time judgment of speed and load change trends. The anomaly judgment module maps different threshold trigger types to event levels, which can be further subdivided into operational warning levels and emergency pause levels. Corresponding to the event level, the scheduling intervention control module selectively executes a deceleration interpolation algorithm or a safety interruption algorithm, directly adjusting the current motion control output or issuing an immediate stop command to the lower-level hardware controller. It can also be expanded according to business needs to implement additional safety measures such as voice prompts, light signal alarms, or system data logging, adapting to the safety response needs of different application scenarios.

[0070] Example: In the field of autonomous decision-making for intelligent agents, when a collaborative robot is running on a production line, if it detects that the speed data momentarily exceeds the upper limit of the process requirements, it will trigger a deceleration intervention to prevent impact on products or workers.

[0071] In the healthcare sector, nursing robots, while providing bedside services to patients, monitor environmental distance data to ensure it is below a safe distance threshold. If this distance is less than the threshold, the robot will promptly stop moving and replan its path to prevent collisions with patients or medical equipment.

[0072] In the fintech business, the service robots in the branch office monitor abnormal load fluctuation data (such as external pulling behavior) while guiding customers, automatically slow down and issue safety prompt voices to ensure the safety and controllability of the service process.

[0073] This embodiment enables the robot system to promptly detect behaviors that deviate from the normal safety range and make targeted adjustments or stops by real-time monitoring and hierarchical intervention of the state data of the controlled object during execution. This ensures the dynamic consistency between the robot's behavior and environmental safety, and significantly reduces safety risks and the uncertainty of system operation.

[0074] S60, collect the execution process data of the security action strategy, and update the strategy generation module based on the execution process data.

[0075] In this embodiment, the execution process data of the safety action strategy refers to the set of various parameters related to the strategy execution status recorded during actual operation. This includes initial action strategies, safety action strategies, actual action sequences, abnormal intervention records, and task completion time data. These data reflect the complete process between strategy execution, safety projection, and real-time intervention. Data collection is achieved through a runtime data recording module, which stores the aforementioned data sources in an ordered manner using timestamps as indexes. Updating refers to dynamically adjusting the strategy generation module based on analysis results, enabling it to generate better initial action strategies in subsequent tasks. Analysis includes converting the difference between the actual action sequence and the safety action strategy into quantifiable action parameter correction frequencies. These frequencies can be further subdivided into joint angle correction rates and path adjustment frequencies, used to measure the intensity of safety projection adjustments. Task completion time data is used to calculate strategy efficiency indicators, measuring the strategy's performance in task timeliness. Abnormal intervention records are used to statistically analyze the frequency of abnormal events, intervention types, and triggering conditions, thereby quantifying the reliability indicators of the safety strategy. By comprehensively considering reliability indicators, strategy efficiency indicators, and action parameter correction frequencies, a safety effectiveness quantification value and a task efficiency balance coefficient are generated. These two serve as co-optimization parameters to simultaneously optimize safety performance and efficiency. The updated policy generation module includes adjusting reward function parameters to increase the penalty weight for safety events, adjusting the task efficiency reward factor, and optimizing the tolerance threshold for motion range limitations during safety projection, making it more accurately constrain the feasible solution space of future action policies. The final update is integrated into the policy generation module through an incremental learning algorithm, ensuring gradual adaptation to environmental and task changes during operation.

[0076] By deploying a data acquisition interface module, the system records in real time the input actions, adjusted actions, execution sequences, intervention times, intervention types, and task completion times for each policy execution. A difference calculation module analyzes the alignment between the actual action sequences and the safety action strategy frame-by-frame, generating joint angle correction rates and path adjustment frequencies as action parameter correction frequency indicators. A statistics module standardizes the task completion time data and calculates the average task efficiency. An abnormal intervention record analysis module categorizes intervention records by event type and calculates the reliability index of abnormal events. Based on these analysis results, an optimizer module uniformly generates safety efficiency quantification values ​​and task efficiency balance coefficients, which are used to adjust the weight parameters related to safety and efficiency in the reward function of the policy generation module. When adjusting the tolerance threshold of the safety projection, the positional constraints of high correction frequencies are tightened, while areas with low correction frequencies are allowed more freedom of action, forming a dynamic adjustment mechanism. Updated parameters are passed to the policy generation module through an incremental learning module, allowing the module parameters to be adjusted gradually, avoiding performance fluctuations caused by global retraining. The data acquisition dimensions can also be expanded according to different task requirements, such as introducing environmental context data and user interaction feedback data, to improve the generalization ability of the policy generation module.

[0077] Example: In the field of autonomous decision-making for intelligent agents, factory mobile robots adjust their strategy generation modules based on execution data analysis across multiple work shifts, thereby reducing the frequency of path corrections and improving stability when encountering complex working conditions.

[0078] In the field of healthcare, surgical robots collect data on the execution process and records of abnormal interventions to gradually optimize the safety weights of the motion strategy generation module, reduce intervention events caused by posture deviations, and ensure the accuracy and safety of medical operations.

[0079] In the fintech business, service robots adjust efficiency balance coefficients and action correction parameters based on collected task completion time and abnormal trigger data, making them both efficient and environmentally adaptable during customer guidance, reducing path adjustments and improving customer experience.

[0080] This embodiment can dynamically optimize the strategy generation module by collecting and analyzing the execution process data of the safety action strategy, making it more adaptable to the current environment and task characteristics, improving the safety and efficiency of the action strategy, and reducing the frequency of adjustments and the probability of safety events during operation.

[0081] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as autonomous decision-making by intelligent agents, fintech, and healthcare. It discloses a method, apparatus, device, and medium for enhancing the safety of action strategies, comprising: constructing a safety constraint space; acquiring multimodal perception data to generate fused safety perception features; generating an initial action strategy based on the task objective and the fused safety perception features through a strategy generation module; projecting the initial action strategy onto the safety constraint space to obtain a safe action strategy; monitoring the state data of the controlled object during the execution of the safe action strategy and implementing intervention measures when the state data triggers a monitoring threshold; collecting execution process data of the safe action strategy and updating the strategy generation module based on the execution process data. This invention, by fusing multimodal perception information with the task objective to generate an initial action strategy and projecting the initial action strategy onto a safety constraint space containing physical constraints, environmental constraints, and preset conditions, and by combining state data monitoring and execution process data feedback to update the strategy generation module, can dynamically ensure the safety of robots performing tasks in complex environments.

[0082] In one embodiment, step S10 above includes:

[0083] S101, Collect the mechanical structure parameters and motion characteristic parameters of the controlled object;

[0084] S102, Based on the mechanical structure parameters and motion characteristic parameters, generate physical safety constraints including motion range limits, load capacity limits and movement speed limits;

[0085] S103, Receive environmental perception data collected by the sensor, and extract spatial distribution information from the environmental perception data;

[0086] S104, Based on the spatial distribution information, generate environmental safety constraints including obstacle distribution information and hazardous area location information;

[0087] S105, access the preset safety conditions library to obtain preset safety conditions including safety distance requirements and prohibitions on dangerous behaviors;

[0088] S106, the physical security constraints, environmental security constraints and preset security conditions are encoded into a mathematical constraint set to generate a security constraint space.

[0089] In this embodiment, firstly, the mechanical structure parameters and motion characteristic parameters of the controlled object are collected. The mechanical structure parameters may include, but are not limited to, the number of joints, the joint arrangement order, the link length, the maximum working radius of the motion arm, the moment of inertia, and the physical shape description of each component. These parameters are derived from the design documents of the controlled object or obtained through 3D scanning and modeling tools. The motion characteristic parameters may include acceleration limits, moment of inertia, maximum load capacity, maximum travel speed, and dynamic equilibrium parameters. These data can be obtained through experimental calibration or analysis of historical operating data. The mechanical structure parameters and motion characteristic parameters serve as basic inputs to calculate the achievable working range of the controlled object in space, establish joint motion range limits, define the upper limit of the load capacity of each component through load limit data, and set speed limit conditions in conjunction with the maximum operating speed, thereby generating physical safety constraints.

[0090] Secondly, the system receives environmental perception data collected by sensors, which may come from LiDAR, depth cameras, infrared sensors, ultrasonic ranging modules, etc. The extraction of spatial distribution information involves converting multi-source data into a dense point cloud in a unified coordinate system, and then using clustering and segmentation algorithms to extract spatial markers representing obstacles and area boundaries. Based on this spatial distribution information, the system calculates the three-dimensional position coordinates and dimensions of obstacles in space, and simultaneously marks the position boundaries of dynamic or static hazardous areas. This generates environmental safety constraints, ensuring that the controlled object can identify and avoid potential environmental risks during movement.

[0091] Furthermore, access to the preset safety condition library is available. This library contains safety distance requirements and prohibited hazardous behaviors. Safety distance requirements can be defined as the minimum safe distance between humans and robots, or between robots and objects. Prohibited hazardous behaviors can include prohibitions on entering specific areas, prohibitions on high-frequency vibration operations, and prohibitions on unauthorized actions. The safety condition library data can come from industry standards, safety regulations, or user-defined configurations. Access is achieved by retrieving and calling its contents through the interface module.

[0092] Finally, physical safety constraints, environmental safety constraints, and preset safety conditions are transformed into a unified constraint set encoding through formal mathematical expressions. This mathematical constraint set can be expressed using matrix inequalities, constraint region sets, or combinations of linear and nonlinear functions, ensuring that all constraints have a clear computational form and a unified interface that can be called by the strategy generation module. Through this encoding, the physical attribute limitations of the controlled object, environmental risk information, and management strategies are integrated to form a complete safety constraint space, providing strict boundary conditions for subsequent action strategy projection and constraint verification.

[0093] This embodiment constructs a complete safety constraint space, which can simultaneously cover the physical limitations of the controlled object, environmental risk information, and preset safety management requirements. This ensures that the generation and optimization of action strategies are always carried out within a strict multi-dimensional constraint range, thereby improving the operational safety and reliability of the controlled object in complex dynamic environments, reducing unnecessary safety risks, and improving the availability of the system.

[0094] In one embodiment, step S20 above includes:

[0095] S201, collects multimodal perception data including visual data, audio data, spatial distance data, and status data;

[0096] S202, Perform object detection on the visual data, extract object recognition features including object contour and texture features based on the detection results, and generate visual features containing object recognition features.

[0097] S203, perform spectrum conversion on the audio data, extract voiceprint features including frequency distribution characteristics from the conversion result, and generate sound features containing voiceprint features;

[0098] S204, perform point cloud reconstruction on the spatial distance data, generate three-dimensional topological features including spatial topological relationships based on the reconstruction results, and generate spatial structural features containing three-dimensional topological features.

[0099] S205, normalize the state data, generate motion parameters including position and velocity parameters based on the processing result, and generate normalized state data containing motion parameters.

[0100] S206, based on the visual features, the sound features, the spatial structure features, and the normalized state data, feature association analysis is performed through a multi-head self-attention mechanism to generate fused security perception features including cross-modal association features.

[0101] In this embodiment, visual data, audio data, spatial distance data, and state data are first acquired. These data form the components of multimodal perception data. Visual data can be captured by a camera module, capturing a sequence of two-dimensional image frames or a video stream. Audio data is acquired by a microphone array, collecting ambient sound signals. Spatial distance data is obtained by measuring the distance distribution of spatial points using a LiDAR or depth camera. State data comes from real-time motion parameters such as position and velocity recorded by motion sensors and a built-in encoder. All perception data must be aligned in a unified timestamp and coordinate system to ensure consistency in subsequent processing.

[0102] Visual data processing employs object detection algorithms, such as convolutional neural network-based detectors, to locate and define target regions in images. Object contour features are extracted from image pixel gradients using edge detection algorithms, while texture features are extracted from local texture patterns using filter banks. These two methods are combined to form object recognition features, ultimately organized into visual features describing the target's shape and surface properties. Audio data is converted into a spectrogram after short-time Fourier transform. The spectrogram transformation reveals the signal's frequency distribution, and voiceprint features are extracted using parameters such as Mel-frequency cepstral coefficients to represent sound source characteristics, organized into sound features used to distinguish sound sources or patterns. Spatial distance data is transformed into a three-dimensional spatial point set through point cloud reconstruction. Spatial clustering and surface fitting algorithms are used to analyze the spatial relationships between points, obtaining spatial topological relationships, organized into spatial structural features to describe the geometric relationships and distribution of objects in the environment. State data undergoes normalization processing, converting its value range into a dimensionless unified interval through linear scaling or normalization methods. After normalization, it can be directly fused with other modal features. After processing, position and velocity parameters are extracted to form motion parameters, which are then organized into normalized state data.

[0103] Multimodal feature fusion employs a multi-head self-attention mechanism. Self-attention computation models and strengthens the relationships between different features by assigning learnable weights to each set of feature vectors. Visual features, auditory features, spatial structure features, and normalized state data interact within this mechanism. Each attention head independently captures the correlations between different modalities or different time scales. Finally, the outputs of all attention heads are integrated through concatenation and linear mapping to generate a fused safety-aware feature containing cross-modal correlation features. This feature serves as the core input for generating safety-enhancing action strategies.

[0104] This embodiment, by uniformly collecting and processing multimodal perception data and combining it with a self-attention mechanism to establish cross-modal correlation analysis, can systematically integrate visual, audio, spatial and motion information, enhance the perception layer's ability to characterize the environment and state in multiple dimensions, and enable subsequent action strategies to take into account environmental changes, object dynamics and behavioral context when making decisions, thereby improving overall perception accuracy and decision robustness and effectively reducing potential safety risks.

[0105] In one embodiment, step S30 above includes:

[0106] S301, parse the task objective to generate a task description including task type and task parameters;

[0107] S302, the task description is transformed into a task feature representation including task priority and key operation points;

[0108] S303, the task feature representation and the fused security perception feature are dimensionally aligned and concatenated to generate a comprehensive input feature that includes task semantics and environmental security status;

[0109] S304, The comprehensive input features are processed by the feature extraction layer of the reinforcement learning policy network of the policy generation module to generate a feature representation including task environment-related features;

[0110] S305, The feature representation is processed through the fully connected layer of the reinforcement learning policy network to generate an action space mapping including the probability distribution of action types;

[0111] S306, The action space mapping is processed through the output layer of the reinforcement learning policy network to generate a time-series action instruction including continuous time action parameters;

[0112] S307, encapsulate the time-series action instructions and generate an initial action strategy.

[0113] In this embodiment, during the parsing of the task objective, the input task objective data is transformed into a structured task description. The task objective data can be natural language text or a predefined instruction set, and its content includes task types such as transportation, inspection, and service, as well as task parameters such as location coordinates, target object category, and priority requirements. When parsing the task type from the task objective data, a classifier or rule mapping method is required. The parsing of task parameters requires format validation and unit conversion to ensure suitability for subsequent processing. The task description, as structured data, contains the task type and task parameters, providing clear semantic boundaries for subsequent transformations.

[0114] When transforming the task description into a task feature representation, a task encoder is used to embed the content of the task description into the feature space. Task priority reflects the importance of the task through weighted encoding, and key operation points are converted into feature vectors through step extraction or predefined task flow templates. These task features are standardized to a uniform dimension to ensure the feasibility of subsequent concatenation operations.

[0115] Aligning and stitching task feature representations with fused security-aware features in terms of dimensions is crucial for achieving multi-source heterogeneous data fusion. By adjusting tensor expansion, padding, or linear transformation, both are brought to the same dimensional structure. The stitching operation generates comprehensive input features using dimensional concatenation or channel stacking. These comprehensive input features integrate information from task semantics and environmental security status, serving as input to the policy generation module.

[0116] The feature extraction layer of the reinforcement learning policy network in the policy generation module abstracts the comprehensive input features through a combination of convolutional units, normalization layers, and activation functions, extracting multi-level, high-dimensional task-environment-related features that reflect the combination of task requirements and the current environment. When processing feature representations, the fully connected layer maps high-dimensional features to the action space. Through a series of weight matrices and bias terms, it generates a probability distribution of action types. Action types include basic robot operations such as movement, rotation, grasping, and placement. The probability of each action type represents the relative likelihood of performing that action.

[0117] When processing motion space mapping in the output layer, softmax normalization and subsequent motion parameter decoding are used to transform discrete motion type probability mappings into continuous time motion parameters, such as velocity vectors, angular velocities, and joint angle sequences, generating directly executable time-series motion instructions. The encapsulation operation, through data format conversion, timestamp synchronization, and motion sequence sorting, organizes the time-series motion instructions into a complete data structure, which is then output as the initial motion strategy to downstream modules.

[0118] This embodiment enables the robot to fully consider task requirements and environmental safety status before performing actions by parsing structured task descriptions, standardizing and encoding task features, and aligning and stitching together dimensions that integrate safety perception features. By combining the hierarchical processing and abstract expression capabilities of reinforcement learning policy networks, the robot achieves a complete link mapping from task semantics to physical execution parameters for action strategies. This improves the rationality, adaptability, and environmental safety of action strategies, providing strong support for robots to perform complex tasks in dynamic environments.

[0119] In one embodiment, step S40 above includes:

[0120] S401, Extract the motion execution parameters, including joint angle parameters, movement speed parameters, and load requirement parameters, from the initial motion strategy;

[0121] S402, Obtain physical safety constraint benchmarks including movement range limits, movement speed limits, and load capacity limits from the safety constraint space;

[0122] S403, compare the joint angle parameter with the range of motion limit, the movement speed parameter with the movement speed limit, and the load requirement parameter with the load capacity limit;

[0123] S404, when any of the joint angle parameter, movement speed parameter and load requirement parameter exceeds the corresponding limit, the joint angle parameter, movement speed parameter or load requirement parameter that exceeds the limit is marked as a violation action parameter;

[0124] S405, adjust the boundary constraint parameters of the violation action parameters to generate corrected action parameters including corrected joint angle, corrected movement speed or corrected load requirement;

[0125] S406, Based on the environmental safety constraints obtained from the safety constraint space, verify the environmental feasibility of the modified action parameters;

[0126] S407, when the environmental feasibility verification is passed, the violation parameters in the initial action strategy are replaced with the corrected action parameters to generate a safe action sequence including timing action instructions;

[0127] S408, encapsulate the security action sequence to generate a security action policy.

[0128] In this embodiment, the motion execution parameters in the initial motion strategy are the specific expression of the machine's motion intention, including joint angle parameters, movement speed parameters, and load requirement parameters. Joint angle parameters represent the joint states of multi-degree-of-freedom components such as the robotic arm, defining the position and posture of each joint using angle values. Movement speed parameters are quantitative indicators of linear or angular velocity during execution, used to describe the robot's required movement speed when performing a task. Load requirement parameters define the weight and torque that the task needs to bear or move, and are one of the fundamental data for safe robot operation. Extracting motion execution parameters requires traversing the motion nodes or sequences in the initial motion strategy, extracting the corresponding parameter value set by timestamp or motion unit.

[0129] The physical safety constraint benchmarks obtained from the safety constraint space serve as baselines for determining whether motion parameters conform to the safety range. Range of motion limits define the permissible angle range of mechanical joints, speed limits constrain the robot's maximum safe operating speed, and load capacity limits the maximum mass the robot can bear under different working conditions. Safety constraint benchmarks are obtained through lookup tables or indexed data structures and are typically stored in configuration files or databases within the robot control system.

[0130] When comparing motion execution parameters with safety constraint benchmarks, the joint angle parameters are compared with the range of motion limit, the movement speed parameters with the movement speed limit, and the load requirement parameters with the load capacity limit. If any parameter is found to exceed the safety threshold, the parameter is marked as a violation motion parameter for subsequent correction processing.

[0131] Adjusting boundary constraint parameters for non-compliant motions is a core measure to ensure that motion execution parameters remain within safe limits. The adjustment process involves trimming or scaling operations to compress joint angles exceeding the limits to within the range of motion, reduce excessive movement speeds to within the movement speed limits, and adjust excessive load demands to below the load capacity limit. The corrected joint angles, corrected movement speeds, and corrected load requirements together constitute the corrected motion parameters, used to replace the original parameters in the initial motion strategy that posed safety risks.

[0132] After correction, the environmental feasibility of the corrected action parameters needs to be verified based on the environmental safety constraints within the safety constraint space. When verifying the environmental feasibility of the corrected action parameters, the specific content of the environmental safety constraints within the safety constraint space must first be clarified. This content includes obstacle distribution information, hazardous area location information, and dynamic obstacle prediction areas. This information originates from real-time sensor data and a pre-set dataset from the environmental model, forming a global spatial safety map. During verification, the motion trajectory corresponding to the corrected action parameters needs to be simulated and mapped onto this spatial safety map. Using a temporal position projection method, it is gradually checked whether the robot will enter an obstacle or hazardous area at any time or spatial point after the corrected action parameters are executed. For example, in a healthcare scenario, when the robot performs a corrected joint angle adjustment, the system needs to verify whether the joint angle will cause the robotic arm's end effector to collide with the bed or medical equipment; in a financial service scenario, it is necessary to verify whether the corrected movement speed and trajectory will cause the robot to enter densely populated areas of the customer service area, especially unauthorized areas. At the algorithm verification level, spatial occupancy grid detection can be used to compare the predicted trajectory generated by the corrected action parameters with the boundary grid of the environmental safety constraints one by one. If the trajectory is found to overlap with the grid of obstacles or danger zones, the environmental feasibility is deemed unsuccessful. If all trajectory nodes are within the safe area, the corrected action parameters are confirmed to meet the environmental feasibility requirements, and the process is allowed to proceed to the subsequent action strategy generation stage.

[0133] After successful verification, the corresponding non-compliant action parameters in the initial action strategy are replaced with corrected action parameters to form a safe action sequence. The safe action sequence consists of action units organized in chronological order, ensuring that the actions executed at each point in time are within physical and environmental safety constraints. Finally, the safe action sequence is encapsulated to generate a safe action strategy. The encapsulation process includes timestamp synchronization, parameter consistency checks, and data structure standardization, giving the safe action strategy completeness that allows it to be directly issued to the robot for execution.

[0134] This embodiment verifies and corrects each safety constraint of the initial action strategy's action execution parameters to ensure that the generated safe action strategy can be executed safely and reliably under physical capabilities, safe speed, load-bearing capacity, and environmental safety conditions. This significantly reduces the risk of collisions, loss of control, and dangerous behaviors caused by unsafe action commands in robot tasks, thereby enhancing the safety of the entire action decision-making process from the source.

[0135] In one embodiment, step S50 above includes:

[0136] S501 collects real-time status data, including position data, speed data, attitude data, and environmental distance data, during the execution of safety action strategies;

[0137] S502, obtains monitoring thresholds including safety distance threshold, movement speed threshold and load fluctuation threshold from the safety monitoring configuration;

[0138] S503, compare the environmental distance data with the safe distance threshold, and when the environmental distance data is less than the safe distance threshold, generate an obstacle approach event;

[0139] S504, compare the speed data with the motion speed threshold, and when the speed data is greater than the motion speed threshold, generate an overspeed motion event;

[0140] S505, monitor the deviation between the load fluctuation data and the load fluctuation threshold, and generate a load anomaly event when the deviation exceeds the load fluctuation threshold;

[0141] S506, Based on the event type of the obstacle approach event, overspeed movement event, or load abnormality event, determine the abnormal event level, including operation warning level and emergency pause level;

[0142] S507, When the abnormal event level is an operation warning level, deceleration intervention measures including adjustment of action parameters are executed;

[0143] S508, when the abnormal event level is emergency pause level, execute interruption intervention measures including stopping movement and initiating safety protocols;

[0144] S509 records the execution log, including the event type and intervention measures.

[0145] In this embodiment, during the execution of the safety action strategy, real-time status data of the controlled object is first collected, including position data, velocity data, attitude data, and environmental distance data. Position data is obtained through a high-precision positioning module, such as a combination of lidar and inertial navigation systems, recording the three-dimensional spatial coordinates of the controlled object. Velocity data is calculated through the fusion of encoder and motor control feedback to provide real-time motion speed information. Attitude data is calculated using a multi-axis gyroscope and accelerometer to obtain precise tilt angles, yaw angles, and roll angles. Environmental distance data is collected by a distance sensing module, such as using an ultrasonic rangefinder, infrared sensor, and 3D depth camera, to continuously and dynamically form measurement results of surrounding obstacles and spatial boundaries. All collected data constitutes a real-time status data set.

[0146] The monitoring module obtains a set of monitoring threshold parameters from the safety monitoring configuration, including a safety distance threshold, a movement speed threshold, and a load fluctuation threshold. The safety distance threshold limits the minimum safe distance between the controlled object and potential obstacles in the environment, ensuring a spatial safety margin. The movement speed threshold limits the maximum safe movement speed of the controlled object, preventing potential hazards caused by high-speed movement. The load fluctuation threshold determines the maximum permissible range of load variation to avoid structural or operational risks caused by abnormal loads.

[0147] The real-time comparison module compares environmental distance data with a safe distance threshold. If any measured distance is lower than the safe distance threshold, an obstacle approach event is generated. It compares speed data with a motion speed threshold; if the current speed exceeds the threshold, an overspeed event is generated. It compares load fluctuation data with a load fluctuation threshold, calculating the deviation between the current load and the baseline load. If the deviation exceeds the threshold, a load anomaly event is generated. These comparison and calculation operations are implemented using independent threshold judgment functions, ensuring real-time performance and calculation accuracy.

[0148] After an event is generated, the system uses the abnormal event level determination module to determine the abnormal event level based on the event type and a predefined risk level configuration table. If the event level is determined to be an operational warning level, the control parameters are immediately adjusted. The speed command output by the motion controller is reduced through the control interface, and the acceleration parameters are adjusted to slow down the dynamic behavior of the controlled object. If the event level is an emergency stop level, a stop command is immediately issued through the safety control interface to stop the operation of the controlled object's drive unit. At the same time, the mechanical locking device or electronic braking mechanism is activated to ensure that the controlled object completely stops moving. Emergency stop level intervention measures also include initiating safety protocol procedures, such as sending an alarm signal to the safety management system through the communication interface and triggering safety lighting or audible and visual alarm devices.

[0149] Throughout all interventions, the system comprehensively records event data. Execution logs include the event type, trigger threshold, a snapshot of the status data at the trigger time, and the intervention measures taken. The log data is structured and stored in the internal security data module for subsequent security analysis, strategy optimization, and dynamic adjustment of security monitoring thresholds. This recording mechanism ensures traceability and data support for all abnormal states and intervention responses, providing accurate data for subsequent system optimization and security performance improvement.

[0150] This embodiment collects and analyzes multi-dimensional state data, including position, speed, attitude, and environmental distance, in real time during the execution of safety action strategies. This data is then dynamically compared with preset safety distance thresholds, movement speed thresholds, and load fluctuation thresholds in the safety monitoring configuration. This allows for the timely detection of controlled objects approaching obstacles, exceeding speed limits, or exhibiting abnormal loads. The system then differentiates and handles these abnormal events according to predefined levels, implementing refined deceleration or interruption interventions. This significantly improves the reliability of controlled objects' safe operation in dynamic environments and reduces the probability of collisions, misoperations, and structural risks. Simultaneously, log recording creates a complete anomaly handling data chain, providing quantifiable and traceable data support for subsequent safety monitoring threshold optimization and strategy generation module updates, achieving closed-loop optimization of safety performance based on behavioral feedback.

[0151] In one embodiment, step S60 above includes:

[0152] S601 collects execution process data including initial action strategy, safety action strategy, actual executed action sequence, abnormal intervention record and task completion time data;

[0153] S602, Analyze the difference between the actual executed action sequence and the safety action strategy, and determine the action parameter correction frequency, including the joint angle correction rate and the path adjustment frequency;

[0154] S603, determine the strategy efficiency index based on the task completion time data;

[0155] S604, Determine the reliability index of the security strategy based on the abnormal intervention record;

[0156] S605, Based on the reliability index, strategy efficiency index and action parameter correction frequency, generate collaborative optimization parameters including safety efficiency quantification value and efficiency balance coefficient;

[0157] S606, adjust the safety weight of the reward function according to the safety efficiency quantification value, and adjust the task efficiency weight of the reward function according to the efficiency balance coefficient to generate optimized reward function parameters;

[0158] S607, Adjust the tolerance threshold of the motion range limitation according to the motion parameter correction frequency, and generate optimized boundary constraint parameters;

[0159] S608, Integrate the optimized reward function parameters and optimized boundary constraint parameters to generate update parameters for the strategy generation module;

[0160] S609, The update parameters of the policy generation module are integrated into the policy generation module through the incremental learning algorithm to generate the updated policy generation module;

[0161] S610, Analyze the event type distribution characteristics of the abnormal intervention records, and update the monitoring threshold in the security monitoring configuration based on the event type distribution characteristics.

[0162] In this embodiment, execution process data is collected, including initial action strategy, safety action strategy, actual action sequence, anomaly intervention record, and task completion time data, to ensure multi-dimensional recording of the robot's entire operation in a real task environment. The initial action strategy represents the original control plan generated by the strategy generation module; the safety action strategy represents the execution instructions corrected by safety constraint space projection; the actual action sequence records the robot's specific operations during the execution phase; the anomaly intervention record includes intervention behaviors executed due to state data triggering monitoring thresholds; and the task completion time data records the time consumed from task start to completion. All data is structured and tagged through a unified data acquisition interface, supporting subsequent data analysis and correlation calculations.

[0163] The differences between the actual executed motion sequence and the safety motion strategy are analyzed to determine the motion parameter correction frequency, including joint angle correction rate and path adjustment frequency. The joint angle correction rate is calculated by comparing the motion trajectory of each joint in the actual executed motion with the joint angle reference value in the corresponding safety motion strategy, and calculating the deviation frequency and amplitude distribution to measure motion consistency at the microscopic level. The path adjustment frequency is analyzed by examining the deviation between the robot's global motion path and the expected path of the safety motion strategy, and statistically analyzing the number of corrections and the length of correction segments to reflect the adjustment trend of the macroscopic path.

[0164] Based on task completion time data, strategy efficiency indicators are determined by comparing actual task completion times with preset task time requirements to form a time efficiency ratio, which describes the strategy execution efficiency and serves as an important input for adjusting the reward function. Based on anomaly intervention records, reliability indicators for the security strategy are determined by statistically analyzing the frequency, triggering conditions, and intervention levels of various anomalies to form a stability metric for secure operation, serving as a parameter basis for measuring the strategy's security performance in different environments.

[0165] Based on reliability indicators, strategy efficiency indicators, and action parameter correction frequency, collaborative optimization parameters are generated, including a safety effectiveness quantification value and an efficiency balance coefficient. The safety effectiveness quantification value is jointly calculated by the reliability indicators and the action parameter correction frequency, reflecting the overall safety performance of the strategy. The efficiency balance coefficient adjusts the safety effectiveness weights through a weighted adjustment of the strategy efficiency indicators, balancing the trade-offs between safety and efficiency to ensure that the optimization result achieves a reasonable balance between safety and task timeliness.

[0166] The safety weights of the reward function are adjusted based on the quantified safety effectiveness values ​​to enhance its sensitivity to safety constraints, enabling the policy generation module to strengthen its preference for safety objectives in subsequent learning. The task efficiency weights of the reward function are adjusted based on the efficiency balance coefficient, ensuring the reward function also reflects a preference for efficient task completion, resulting in optimized reward function parameters. The tolerance threshold for motion range constraints is adjusted based on the motion parameter correction frequency, dynamically adjusting the robot's motion space constraints. When the correction frequency is high, the motion space is automatically tightened to reduce potential risk space and improve safety, resulting in optimized boundary constraint parameters.

[0167] The optimized reward function parameters and optimized boundary constraint parameters are integrated to generate update parameters for the policy generation module, serving as a unified input for enhancing the module's functionality. An incremental learning algorithm is used to integrate these updated parameters into the policy generation module, ensuring that the latest data-driven optimization results are incorporated without losing historical training outcomes. This results in a more secure and adaptable policy generation module.

[0168] Analyze the event type distribution characteristics of abnormal intervention records, and statistically analyze the occurrence frequency and distribution patterns of various events in different time periods and operating environments. Based on these characteristics, dynamically adjust various monitoring thresholds in the security monitoring configuration to improve the sensitivity and scene adaptability of the monitoring strategy and support the system's rapid response to complex environmental changes.

[0169] Example: In the field of autonomous decision-making for intelligent agents, for example, when an autonomous mobile service robot is moving sensitive items in an indoor environment, it completes the task based on a complete action strategy and safety enhancement process.

[0170] First, the agent constructs a safety constraint space based on the current task instructions. The system collects the robot's own mechanical structural parameters, such as maximum joint angles, maximum load capacity, and maximum speed, as well as kinematic model parameters, to generate physical safety constraints, including limits on range of motion, load capacity, and speed. Simultaneously, the robot uses LiDAR and depth cameras to acquire environmental perception data. Through the spatial distribution and structural reconstruction of objects in the environment, it determines obstacle distribution information and the location of hazardous areas, forming environmental safety constraints. Furthermore, the robot accesses preset safety conditions from the task configuration database, including personnel safety distances and operational restrictions, encoding these into a unified set of mathematical constraints. This ultimately forms the safety constraint space, serving as the boundary conditions for subsequent strategy decisions.

[0171] After the constrained space is constructed, the intelligent robot collects multimodal perception data, including visual data (identifying people, objects, and landmarks), audio data (detecting environmental noise and abnormal sounds), spatial distance data (assessing distances to surrounding obstacles), and state data (position, posture, and velocity). Visual data undergoes object detection and contour and texture feature extraction to generate visual features; audio data is transformed through spectrum conversion to obtain voiceprint features, generating sound features; spatial distance data is reconstructed from point clouds to generate three-dimensional topological features, forming spatial structure features; and state data is normalized to provide standardized position and velocity parameters. All these features are analyzed across modalities using a multi-head self-attention mechanism to generate fused safety perception features that reflect the multidimensional dynamic state of the robot's current task environment.

[0172] Based on the target task requirements (e.g., safely moving an item from a designated starting point to a target area), the task type (movement task) and task parameters (target location, required time, etc.) are parsed and transformed into task feature representations of task priority and key operation points. This task feature representation, after being dimensionally aligned and concatenated with the fused safety awareness features, is input into the reinforcement learning policy network of the policy generation module. After processing by the feature extraction layer, task environment-related features are obtained, which are then mapped to action type probability distributions through a fully connected layer. Finally, the output layer decodes and generates time-series action instructions with continuous-time action parameters, which are then encapsulated to form the initial action policy.

[0173] After the initial motion strategy is formed, the robot projects it onto the constructed safety constraint space. At this point, the robot extracts motion execution parameters and compares them one by one with joint angles and range of motion limits, movement speed and speed limits, load requirements and maximum carrying capacity. When any parameter is detected to be outside the allowable range, it is marked as a violation parameter. Boundary constraints are adjusted for the violation parameters to generate corrected joint angles, speeds, and load requirements. The feasibility of these adjustments is further verified based on environmental safety constraints, such as ensuring that the robot's adjusted path does not traverse dynamic crowds or hazardous areas. If the verification passes, the corresponding parameters in the initial strategy are replaced, and the system is repackaged as a safe motion strategy.

[0174] During the execution of the safety action strategy, the robot continuously monitors its status data, collecting its position, speed, posture, and environmental distance from surrounding objects. This data is compared in real-time with the safety distance threshold, movement speed threshold, and load fluctuation threshold set in the safety monitoring configuration. When the environmental distance is below the threshold, an obstacle approach event is generated; when the speed exceeds the threshold, an overspeed event is generated; and when the load fluctuation is abnormal, a load anomaly event is generated. Anomalies are categorized into different levels based on the event type: if it is an operational warning level, the robot automatically adjusts its motion parameters and decelerates; if it is an emergency pause level, it immediately stops moving and activates the safety protocol. All monitored events and executed interventions are recorded in the system log in real-time, forming an execution history.

[0175] After the task is completed, the robot collects and organizes complete execution process data, including the initial action strategy, safety action strategy, actual action sequence, abnormal intervention records, and task completion time data. The system analyzes the deviation between the actual execution and the safety action strategy, calculates the joint angle correction rate and path adjustment frequency as the action parameter correction frequency, calculates the strategy efficiency index in conjunction with the task completion time, and statistically determines the safety strategy reliability index based on the abnormal intervention records. Based on these analysis results, the safety efficiency quantification value and efficiency balance coefficient are calculated as co-optimization parameters, adjusting the safety weight and efficiency weight of the reward function to optimize the reward function parameters; simultaneously, the tolerance threshold of the motion range is dynamically adjusted according to the correction frequency to optimize the boundary constraint parameters. Finally, all optimized parameters are integrated to form the update parameters of the strategy generation module, and integrated into the existing strategy generation module through an incremental learning algorithm to achieve adaptive optimization of the model. The system also dynamically adjusts the monitoring threshold by analyzing the distribution of abnormal event types, enabling the robot to further improve its safety protection level and environmental adaptability in subsequent tasks.

[0176] In the field of healthcare, an autonomous surgical robot assists doctors in performing complex procedures in a dynamic operating room environment. The robot needs to perform its tasks in a highly dynamic, restricted environment with extremely high safety requirements. Every step of its operation must strictly adhere to safety regulations and actively sense and respond to changes in the environment to ensure personal safety and surgical outcomes.

[0177] The robot first defines the boundary conditions for its actions by constructing a safety constraint space. By collecting data on the joint range of motion of its robotic arm, maximum tolerable operating force, and the movement speed of its end effector, the system generates physical safety constraints to prevent operations from exceeding its mechanical limits. Combined with a high-precision depth camera and medical-specific sensors within the operating room, it perceives the spatial distribution information of medical staff, patient positions, and instrument table locations in real time, forming dynamic environmental safety constraints. It then accesses a preset safety condition library in the hospital management system, retrieving conditions such as safety distance requirements, prohibited boundary areas, and emergency stop response requirements. These physical safety constraints, environmental safety constraints, and preset safety conditions are integrated and encoded into a mathematical constraint set, defining the safety constraint space for action execution.

[0178] The robot then collects multimodal perception data to comprehensively understand the current surgical environment. Its vision system captures medical images, instruments, and personnel positions in real time, extracting body contours and texture features for instrument recognition and localization. Audio data is analyzed through spectrum conversion to distinguish between critical surgical instructions and background noise. Spatial distance data is used to reconstruct a 3D surgical scene spatial model through point cloud reconstruction, and state data is normalized to describe the robot's current posture, position, and velocity. This cross-modal data is correlated and analyzed through a multi-head self-attention mechanism to generate fused safety perception features that reflect the real-time dynamic state of the operating room, ensuring that the robot can perceive spatial constraints and personnel dynamics.

[0179] Combined with specific surgical task objectives, such as suturing at specific locations on the organ surface, the robot analyzes the task type (e.g., suturing operation) and task parameters (suturing path, pressure requirements), extracting task priorities (prioritizing accuracy) and key operation points (surgical tip path nodes). These task features are dimensionally aligned and concatenated with fused safety-aware features before being input into a reinforcement learning policy network. The feature extraction layer generates feature representations reflecting the relationship between the task and the environment, which are then passed through a fully connected layer to generate a probability distribution of action types. The output layer forms a series of continuous-time action parameters, which are finally encapsulated into an initial action policy.

[0180] After forming the initial motion strategy, the robot undergoes safety constraint space projection correction, extracting joint angle parameters, movement speed parameters, and load requirement parameters. These parameters are compared with preset motion range limits, movement speed limits, and load capacity limits to detect any non-compliant motion parameters. For parameters exceeding the limits, the system performs boundary constraint adjustments, such as reducing the excessive movement speed to a safe range, ensuring that the force and speed during the suturing process are within a safe and controllable range. The corrected parameters are verified through environmental safety constraints, such as ensuring that the adjusted suturing path will not cross the dynamic path of medical personnel or come into contact with medical instruments. After successful verification, the parameters corresponding to the initial motion strategy are replaced to form the final safe motion strategy.

[0181] During the execution of safety action strategies, the robot collects real-time data on its position, speed, posture, and spatial distance to personnel and instruments in the environment. It monitors these data based on safety monitoring settings, including safe distance thresholds, speed thresholds, and load fluctuation thresholds. When the environmental distance is too small, an obstacle approach event is generated; when the speed is too high, an overspeed event is generated; and when the load fluctuates abnormally, a load anomaly event is generated. Based on the type and severity of the abnormal event, the system intervenes in a tiered manner. Low-risk events trigger deceleration intervention to prevent the scalpel from moving too quickly, while high-risk events immediately trigger an emergency pause, stopping the robotic arm's movements and issuing an alarm to ensure surgical safety. All abnormal events and intervention details are meticulously recorded, forming traceable log data.

[0182] Upon completion of the task, the system collects complete execution process data, including initial action strategy, safety action strategy, actual action sequence, abnormal intervention records, and surgery completion time data. It analyzes the differences between the actual actions and the safety action strategy, determines the correction frequency of action parameters such as joint angle correction rate and path adjustment frequency, calculates task completion time and accuracy as efficiency indicators, and statistically analyzes the frequency of abnormal interventions as a reliability indicator. Based on these results, it generates safety efficiency quantification values ​​and efficiency balance coefficients as co-optimization parameters, updates the safety and efficiency weights of the reward function, and optimizes the parameters. It adjusts the motion range tolerance threshold by combining the action parameter correction frequency, forming updated parameters, and integrates them into the strategy generation module using an incremental learning algorithm to achieve adaptive optimization of the robot to the dynamic environment of the operating room. It updates monitoring thresholds through abnormal event distribution analysis to improve the accuracy and timeliness of subsequent task monitoring and response.

[0183] In the fintech field, an intelligent service robot autonomously assists in handling customer business in a bank branch environment. The robot needs to make safe and autonomous decisions and execute service operations in a complex scenario with limited space and frequent interactions between customers and staff, while ensuring that it does not cause physical risks to customers or affect the normal order of the branch due to uncontrolled actions.

[0184] The system first defines the boundary conditions for the service robot's behavior, constructing a safety constraint space. By collecting the structural parameters of the robot's robotic arm, the acceleration limit of the mobile chassis, the maximum load capacity of the gripping device, and the upper limit of its travel speed, physical safety constraints are established to limit the robot's movement and range of motion in the counter area and customer waiting area. Combined with multimodal sensing devices deployed in the service hall, real-time spatial distribution information such as customer distribution, queue order, counter location, and aisle width is perceived to generate dynamic environmental safety constraints. The system accesses the service hall's operational safety rules database to obtain requirements for service distances, prohibited areas (such as the cashier's operating area), etc., forming preset safety conditions. By integrating physical safety constraints, environmental safety constraints, and preset safety conditions, these are encoded into a mathematical constraint set, defining the robot's global behavioral boundaries.

[0185] The robot then comprehensively perceives the service hall environment through multimodal perception data. The vision system captures customers' facial expressions, gestures, and posture, extracting object recognition features such as contours and textures to identify customer categories (regular customers, elderly customers, etc.) and confirm queue order. Audio data undergoes spectrum analysis to identify customer voice requests and distinguish environmental noise. Spatial distance data is reconstructed using LiDAR point clouds to form a detailed service hall spatial model, including the distance between customers and counters, and aisle accessibility. The robot's own speed and position data are normalized to describe its current motion state. This multimodal information is cross-correlated through a multi-head self-attention mechanism to generate fused safety perception features, reflecting the service hall's dynamics and service needs in real time, ensuring the robot's safe and accurate understanding of the environment.

[0186] Subsequently, the robot, in conjunction with specific service task objectives (such as delivering receipts to customers or guiding customers to the corresponding window), analyzes the task category and service parameters, extracting task priorities (such as prioritizing service for elderly customers) and key operational points (such as the standing service point location and the arm angle for delivering receipts). These task features, along with fused safety perception features, are concatenated through dimensional alignment and input into a reinforcement learning policy network. After feature extraction, fully connected processing, and output layer computation, an action type probability distribution is generated, and finally, continuous-time action parameters are output and encapsulated into an initial action policy.

[0187] The initial motion strategy is further projected onto a safety constraint space for calibration. Motion parameters such as joint angles, movement speed, and load requirements are extracted and compared one by one with the movement range, speed limits, and load limits in the safety constraint space to detect any non-compliant motion parameters. Boundary constraint adjustments are applied to parameters exceeding limits; for example, the travel speed is reduced to below the maximum safe speed allowed in the sales hall passage, and the arm gripping angle is adjusted to ensure that the receipt delivery action is within an acceptable distance for the customer. The adjusted parameters are then verified through environmental safety constraints to ensure that the actions will not cause spatial conflicts with queuing customers or obstruct the main passage. After verification, the robot replaces the corresponding parameters to form a safe motion strategy that includes time-series motion instructions.

[0188] During the execution of safety action strategies, the robot continuously collects data on its position, speed, posture, and distance to the environment, and compares this data in real time with safety distance thresholds, movement speed thresholds, and load fluctuation thresholds configured in the safety monitoring system. When a customer approaches the robot, triggering a "too close" condition, the system records a "customer approach event"; when the robot accelerates excessively, an "overspeed event" is recorded; and when the gripper's load fluctuation exceeds limits, an "abnormal load event" is recorded. Based on the event type and severity, the system performs tiered interventions: low-risk events trigger deceleration adjustments (e.g., slowing down to avoid disturbing the customer), while high-risk events immediately halt operations (e.g., stopping movement immediately and retreating to a safe area). All abnormal events and intervention executions are recorded to generate a complete log, facilitating subsequent business and safety reviews.

[0189] After completing the service task, the system collects data throughout the entire process, including initial action strategies, safety action strategies, actual action sequences, anomaly intervention records, and task completion time data. By analyzing the differences between the actual actions and the safety action strategies, the system calculates the joint angle correction rate and path adjustment frequency, extracting task execution efficiency indicators and service interruption reliability indicators. Based on these analysis results, the system calculates safety efficiency quantification values ​​and efficiency balance coefficients as co-optimization parameters to adjust the safety and efficiency weights of the reward function. These parameters are then combined with the action correction frequency to adjust the motion tolerance threshold, ultimately forming updated parameters that are integrated into the strategy generation module through an incremental learning algorithm. By analyzing the event distribution of anomaly intervention records, the system dynamically optimizes the monitoring thresholds within the service hall, improving the response sensitivity and security level of subsequent operations.

[0190] This embodiment continuously optimizes the reward function and boundary constraint parameters of the strategy generation module by collecting execution process data and combining it with multi-dimensional safety and efficiency index analysis. It achieves online updates through incremental learning algorithms without interfering with robot operation, ensuring that control decisions always conform to the dynamic requirements of the current environment and task. Furthermore, it drives the dynamic adjustment of monitoring thresholds through the distribution of abnormal events, ultimately constructing a complete closed loop from action execution to strategy optimization and then to the linkage of monitoring strategies, significantly improving the safety, adaptability, and stability of robot task execution.

[0191] In one embodiment, a motion strategy security enhancement device is provided, which corresponds one-to-one with the motion strategy security enhancement method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the motion strategy safety enhancement device of the present invention. The modules include a safety constraint modeling module 10, a multimodal perception fusion module 20, a strategy generation module 30, a safety strategy projection module 40, a safety monitoring and intervention module 50, and a strategy adaptive optimization module 60. Detailed descriptions of each functional module are as follows:

[0192] The safety constraint modeling module 10 is used to construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions.

[0193] The multimodal perception fusion module 20 is used to acquire multimodal perception data and generate fused security perception features based on the multimodal perception data.

[0194] The strategy generation module 30 is used to generate an initial action strategy based on the task objective and the fused security awareness features.

[0195] The security policy projection module 40 is used to project the initial action policy onto the security constraint space to obtain a security action policy.

[0196] The safety monitoring and intervention module 50 is used to monitor the status data of the controlled object during the execution of the safety action strategy, and to execute intervention measures when the status data triggers a preset monitoring threshold.

[0197] The policy adaptive optimization module 60 is used to collect the execution process data of the security action policy and update the policy generation module based on the execution process data.

[0198] In one embodiment, the safety constraint modeling module 10 is specifically used for:

[0199] Collect the mechanical structure parameters and motion characteristic parameters of the controlled object;

[0200] Based on the mechanical structure parameters and motion characteristic parameters, physical safety constraints are generated, including motion range limits, load capacity limits, and movement speed limits.

[0201] Receive environmental perception data collected by sensors and extract spatial distribution information from the environmental perception data;

[0202] Based on the spatial distribution information, environmental safety constraints are generated, including obstacle distribution information and hazardous area location information.

[0203] Access the preset safety conditions library to obtain preset safety conditions, including safety distance requirements and prohibitions on dangerous behaviors;

[0204] The physical security constraints, environmental security constraints, and preset security conditions are encoded into a set of mathematical constraints to generate a security constraint space.

[0205] In one embodiment, the multimodal perception fusion module 20 is specifically used for:

[0206] Collect multimodal perception data, including visual data, audio data, spatial distance data, and status data;

[0207] Object detection is performed on the visual data, and object recognition features, including object contour and texture features, are extracted based on the detection results to generate visual features containing object recognition features.

[0208] The audio data is subjected to spectral transformation, and voiceprint features including frequency distribution characteristics are extracted from the transformation result to generate sound features containing voiceprint features;

[0209] Point cloud reconstruction is performed on the spatial distance data, and three-dimensional topological features including spatial topological relationships are generated based on the reconstruction results, thereby generating spatial structural features containing three-dimensional topological features.

[0210] The state data is normalized, and motion parameters including position and velocity parameters are generated based on the processing result, thus generating normalized state data containing motion parameters.

[0211] Based on the visual features, the sound features, the spatial structure features, and the normalized state data, feature association analysis is performed through a multi-head self-attention mechanism to generate fused security perception features that include cross-modal association features.

[0212] In one embodiment, the policy generation module 30 is specifically used for:

[0213] Parse the task objective to generate a task description that includes the task type and task parameters;

[0214] The task description is transformed into a task feature representation that includes task priority and key operation points;

[0215] The task feature representation and the fused security awareness feature are dimensionally aligned and concatenated to generate a comprehensive input feature that includes task semantics and environmental security status.

[0216] The comprehensive input features are processed by the feature extraction layer of the reinforcement learning policy network of the policy generation module to generate a feature representation including task environment-related features;

[0217] The feature representation is processed by the fully connected layer of the reinforcement learning policy network to generate an action space mapping including the probability distribution of action types;

[0218] The action space mapping is processed by the output layer of the reinforcement learning policy network to generate a time-series action instruction including continuous-time action parameters.

[0219] The time-series action instructions are encapsulated to generate an initial action strategy.

[0220] In one embodiment, the security policy projection module 40 is specifically used for:

[0221] Extract the motion execution parameters, including joint angle parameters, movement speed parameters, and load requirement parameters, from the initial motion strategy;

[0222] Obtain physical safety constraint benchmarks, including movement range limits, movement speed limits, and load capacity limits, from the safety constraint space;

[0223] Compare the joint angle parameter with the range of motion limit, the movement speed parameter with the movement speed limit, and the load requirement parameter with the load capacity limit;

[0224] When any of the joint angle parameter, movement speed parameter, and load requirement parameter exceeds the corresponding limit, the joint angle parameter, movement speed parameter, or load requirement parameter that exceeds the limit will be marked as a violation action parameter.

[0225] The boundary constraint parameters of the aforementioned non-compliant action parameters are adjusted to generate corrected action parameters, including corrected joint angles, corrected movement speeds, or corrected load requirements.

[0226] Based on the environmental safety constraints obtained from the safety constraint space, the environmental feasibility of the modified action parameters is verified.

[0227] When the environmental feasibility verification is passed, the violation parameters in the initial action strategy are replaced with the corrected action parameters to generate a safe action sequence including time-sequenced action instructions.

[0228] The security action sequence is encapsulated to generate a security action policy.

[0229] In one embodiment, the security monitoring and intervention module 50 is specifically used for:

[0230] During the execution of safety action strategies, real-time status data, including position data, velocity data, attitude data, and environmental distance data, is collected.

[0231] Obtain monitoring thresholds, including safety distance threshold, movement speed threshold, and load fluctuation threshold, from the security monitoring configuration;

[0232] The environmental distance data is compared with the safe distance threshold. When the environmental distance data is less than the safe distance threshold, an obstacle approach event is generated.

[0233] The speed data is compared with the motion speed threshold. When the speed data is greater than the motion speed threshold, an overspeed motion event is generated.

[0234] Monitor the deviation between load fluctuation data and the load fluctuation threshold. When the deviation exceeds the load fluctuation threshold, generate a load anomaly event.

[0235] Based on the event type of the obstacle approach event, overspeed event, or abnormal load event, determine the abnormal event level, including the operation warning level and the emergency stop level;

[0236] When the abnormal event level is an operational warning level, deceleration intervention measures, including adjustment of action parameters, are implemented.

[0237] When the abnormal event level is emergency pause level, interruption intervention measures including stopping movement and activating safety protocols are implemented;

[0238] Record execution logs including event type and intervention measures.

[0239] In one embodiment, the policy adaptive optimization module 60 is specifically used for:

[0240] Collect execution process data including initial action strategy, safety action strategy, actual action sequence, abnormal intervention records, and task completion time data;

[0241] Analyze the differences between the actual executed action sequence and the safety action strategy to determine the action parameter correction frequency, including joint angle correction rate and path adjustment frequency;

[0242] Determine strategy efficiency indicators based on the task completion time data;

[0243] The reliability indicators of the security strategy are determined based on the aforementioned abnormal intervention records;

[0244] Based on the aforementioned reliability index, strategy efficiency index, and action parameter correction frequency, collaborative optimization parameters including safety efficiency quantification value and efficiency balance coefficient are generated.

[0245] The safety weight of the reward function is adjusted according to the safety efficiency quantification value, and the task efficiency weight of the reward function is adjusted according to the efficiency balance coefficient to generate the optimized reward function parameters;

[0246] Based on the motion parameters, the tolerance threshold for adjusting the motion range limitation is adjusted according to the frequency correction, and optimized boundary constraint parameters are generated.

[0247] Integrate the optimized reward function parameters and optimized boundary constraint parameters to generate update parameters for the strategy generation module;

[0248] The updated parameters of the policy generation module are integrated into the policy generation module through an incremental learning algorithm to generate an updated policy generation module.

[0249] Analyze the event type distribution characteristics of the abnormal intervention records, and update the monitoring thresholds in the security monitoring configuration based on the event type distribution characteristics.

[0250] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a security enhancement method for action policies on the server side.

[0251] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of a security enhancement method on the user side.

[0252] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0253] Construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions;

[0254] Acquire multimodal perception data and generate fused security perception features based on the multimodal perception data;

[0255] Based on the task objective and the fused security awareness features, an initial action policy is generated through the policy generation module.

[0256] The initial action strategy is projected onto the safety constraint space to obtain the safety action strategy;

[0257] During the execution of the safety action strategy, the status data of the controlled object is monitored, and when the status data triggers a preset monitoring threshold, intervention measures are executed.

[0258] Collect execution process data of the security action policy, and update the policy generation module based on the execution process data.

[0259] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0260] Construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions;

[0261] Acquire multimodal perception data and generate fused security perception features based on the multimodal perception data;

[0262] Based on the task objective and the fused security awareness features, an initial action policy is generated through the policy generation module.

[0263] The initial action strategy is projected onto the safety constraint space to obtain the safety action strategy;

[0264] During the execution of the safety action strategy, the status data of the controlled object is monitored, and when the status data triggers a preset monitoring threshold, intervention measures are executed.

[0265] Collect execution process data of the security action policy, and update the policy generation module based on the execution process data.

[0266] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0267] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0268] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0269] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for enhancing the safety of action strategies, characterized in that, Includes the following steps: Construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions; Acquire multimodal perception data and generate fused security perception features based on the multimodal perception data; Based on the task objective and the fused security awareness features, an initial action policy is generated through the policy generation module. The initial action strategy is projected onto the safety constraint space to obtain the safety action strategy; During the execution of the safety action strategy, the status data of the controlled object is monitored, and when the status data triggers a preset monitoring threshold, intervention measures are executed. Collect execution process data of the security action policy, and update the policy generation module based on the execution process data.

2. The motion strategy security enhancement method as described in claim 1, characterized in that, Construct a safety constraint space that includes physical safety constraints based on the physical attributes of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions, including: Collect the mechanical structure parameters and motion characteristic parameters of the controlled object; Based on the mechanical structure parameters and motion characteristic parameters, physical safety constraints are generated, including motion range limits, load capacity limits, and movement speed limits. Receive environmental perception data collected by sensors and extract spatial distribution information from the environmental perception data; Based on the spatial distribution information, environmental safety constraints are generated, including obstacle distribution information and hazardous area location information. Access the preset safety conditions library to obtain preset safety conditions, including safety distance requirements and prohibitions on dangerous behaviors; The physical security constraints, environmental security constraints, and preset security conditions are encoded into a set of mathematical constraints to generate a security constraint space.

3. The motion strategy security enhancement method as described in claim 1, characterized in that, Acquiring multimodal perception data and generating fused security perception features based on the multimodal perception data includes: Collect multimodal perception data, including visual data, audio data, spatial distance data, and status data; Object detection is performed on the visual data, and object recognition features, including object contour and texture features, are extracted based on the detection results to generate visual features containing object recognition features. The audio data is subjected to spectral transformation, and voiceprint features including frequency distribution characteristics are extracted from the transformation result to generate sound features containing voiceprint features; Point cloud reconstruction is performed on the spatial distance data, and three-dimensional topological features including spatial topological relationships are generated based on the reconstruction results, thereby generating spatial structural features containing three-dimensional topological features. The state data is normalized, and motion parameters including position and velocity parameters are generated based on the processing result, thus generating normalized state data containing motion parameters. Based on the visual features, the sound features, the spatial structure features, and the normalized state data, feature association analysis is performed through a multi-head self-attention mechanism to generate fused security perception features that include cross-modal association features.

4. The motion strategy security enhancement method as described in claim 1, characterized in that, Based on the task objective and the fused security awareness features, an initial action policy is generated through the policy generation module, including: Parse the task objective to generate a task description that includes the task type and task parameters; The task description is transformed into a task feature representation that includes task priority and key operation points; The task feature representation and the fused security awareness feature are dimensionally aligned and concatenated to generate a comprehensive input feature that includes task semantics and environmental security status. The comprehensive input features are processed by the feature extraction layer of the reinforcement learning policy network of the policy generation module to generate a feature representation including task environment-related features; The feature representation is processed by the fully connected layer of the reinforcement learning policy network to generate an action space mapping including the probability distribution of action types; The action space mapping is processed by the output layer of the reinforcement learning policy network to generate a time-series action instruction including continuous-time action parameters. The time-series action instructions are encapsulated to generate an initial action strategy.

5. The motion strategy security enhancement method as described in claim 1, characterized in that, Projecting the initial action policy onto the safety constraint space yields a safety action policy, including: Extract the motion execution parameters, including joint angle parameters, movement speed parameters, and load requirement parameters, from the initial motion strategy; Obtain physical safety constraint benchmarks, including movement range limits, movement speed limits, and load capacity limits, from the safety constraint space; Compare the joint angle parameter with the range of motion limit, the movement speed parameter with the movement speed limit, and the load requirement parameter with the load capacity limit; When any of the joint angle parameter, movement speed parameter, and load requirement parameter exceeds the corresponding limit, the joint angle parameter, movement speed parameter, or load requirement parameter that exceeds the limit will be marked as a violation action parameter. The boundary constraint parameters of the aforementioned non-compliant action parameters are adjusted to generate corrected action parameters, including corrected joint angles, corrected movement speeds, or corrected load requirements. Based on the environmental safety constraints obtained from the safety constraint space, the environmental feasibility of the modified action parameters is verified. When the environmental feasibility verification is passed, the violation parameters in the initial action strategy are replaced with the corrected action parameters to generate a safe action sequence including time-sequenced action instructions. The security action sequence is encapsulated to generate a security action policy.

6. The motion strategy security enhancement method as described in claim 1, characterized in that, During the execution of the safety action strategy, the status data of the controlled object is monitored. When the status data triggers a preset monitoring threshold, intervention measures are executed, including: During the execution of safety action strategies, real-time status data, including position data, velocity data, attitude data, and environmental distance data, is collected. Obtain monitoring thresholds, including safety distance threshold, movement speed threshold, and load fluctuation threshold, from the security monitoring configuration; The environmental distance data is compared with the safe distance threshold. When the environmental distance data is less than the safe distance threshold, an obstacle approach event is generated. The speed data is compared with the motion speed threshold. When the speed data is greater than the motion speed threshold, an overspeed motion event is generated. Monitor the deviation between load fluctuation data and the load fluctuation threshold. When the deviation exceeds the load fluctuation threshold, generate a load anomaly event. Based on the event type of the obstacle approach event, overspeed event, or abnormal load event, determine the abnormal event level, including the operation warning level and the emergency stop level; When the abnormal event level is an operational warning level, deceleration intervention measures, including adjustment of action parameters, are implemented. When the abnormal event level is emergency pause level, interruption intervention measures including stopping movement and activating safety protocols are implemented; Record execution logs including event type and intervention measures.

7. The motion strategy security enhancement method as described in claim 1, characterized in that, Collecting execution process data of the security action policy and updating the policy generation module based on the execution process data includes: Collect execution process data including initial action strategy, safety action strategy, actual action sequence, abnormal intervention records, and task completion time data; Analyze the differences between the actual executed action sequence and the safety action strategy to determine the action parameter correction frequency, including joint angle correction rate and path adjustment frequency; Determine strategy efficiency indicators based on the task completion time data; The reliability indicators of the security strategy are determined based on the aforementioned abnormal intervention records; Based on the aforementioned reliability index, strategy efficiency index, and action parameter correction frequency, collaborative optimization parameters including safety efficiency quantification value and efficiency balance coefficient are generated. The safety weight of the reward function is adjusted according to the safety efficiency quantification value, and the task efficiency weight of the reward function is adjusted according to the efficiency balance coefficient to generate the optimized reward function parameters; Based on the motion parameters, the tolerance threshold for adjusting the motion range limitation is adjusted according to the frequency correction, and optimized boundary constraint parameters are generated. Integrate the optimized reward function parameters and optimized boundary constraint parameters to generate update parameters for the strategy generation module; The updated parameters of the policy generation module are integrated into the policy generation module through an incremental learning algorithm to generate an updated policy generation module. Analyze the event type distribution characteristics of the abnormal intervention records, and update the monitoring thresholds in the security monitoring configuration based on the event type distribution characteristics.

8. A motion strategy safety enhancement device, characterized in that, The action strategy security enhancement device includes: The safety constraint modeling module is used to construct a safety constraint space that includes physical safety constraints based on the physical properties of the controlled object, environmental safety constraints based on environmental perception data, and preset safety conditions. A multimodal perception fusion module is used to acquire multimodal perception data and generate fused security perception features based on the multimodal perception data; The strategy generation module is used to generate an initial action strategy based on the task objective and the fused security awareness features. A security policy projection module is used to project the initial action policy onto the security constraint space to obtain a security action policy. The safety monitoring and intervention module is used to monitor the status data of the controlled object during the execution of the safety action strategy, and to execute intervention measures when the status data triggers a preset monitoring threshold. The policy adaptive optimization module is used to collect execution process data of the security action policy and update the policy generation module based on the execution process data.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and an action policy security enhancement program stored in the memory and executable on the processor. When executed by the processor, the action policy security enhancement program implements the steps of the action policy security enhancement method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an action policy security enhancement program, which, when executed by a processor, implements the steps of the action policy security enhancement method as described in any one of claims 1-7.

Citation Information

Cited By

  • Visual correction system and method for biosafety cabinet operation training

    CN121963042A

  • Verifiable execution method and system of dual-system safety shield for fire scene and medium

    CN122034000A

  • Self-checking method, self-checking device and computer equipment

    CN122343455A