Robot safety control method based on multi-modal embodied intelligent agent, electronic device and readable storage medium

The robot safety control method based on multimodal perception and physical coding solves the problems of perception robustness and safety of robots in complex environments, and achieves efficient and reliable safety control.

CN122500699APending Publication Date: 2026-08-04SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST UNIV
Filing Date
2026-05-12
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing robot safety control methods lack robustness in multimodal perception in complex environments and lack the ability to comprehensively reason about task instructions and environmental conditions, resulting in safety risks and low operational efficiency.

Method used

By collecting environmental visual information, voice commands, and robot body state data through multimodal sensors, a key entity encoding mechanism is constructed, multi-stage task reasoning is performed, and safety semantic verification is introduced. Combined with a dynamic safety intervention mechanism, a closed-loop control is formed.

Benefits of technology

It enables safe, continuous, and reliable control of robots in complex human-robot collaborative environments, reduces the risk of safety accidents, and improves operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122500699A_ABST
    Figure CN122500699A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, embodied intelligence and robot safety control, and particularly discloses a robot safety control method based on a multi-modal embodied intelligent agent, an electronic device and a readable storage medium. The method collects multi-modal information of vision, voice, ontology state and human-computer interaction, constructs key entity semantic representation fusing geometric and physical attributes by using cross-modal alignment technology, and realizes deep situation awareness of the intelligent agent to a complex environment. On this basis, the human body key points are subjected to time-space modeling, dynamic action safety semantics are generated, and the dynamic action safety semantics are injected into a causal reasoning process of a multi-stage control instruction as a hard constraint condition, so that the work target is achieved while safety specifications are met. The application constructs a safety closed-loop mechanism of perception, reasoning, execution and feedback, and can monitor environmental mutations in real time and implement dynamic intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, embodied intelligence, and robot safety control technology, and in particular to a robot safety control method, electronic device, and readable storage medium based on a multimodal embodied intelligent agent. Background Technology

[0002] With the development of artificial intelligence and embodied intelligence technologies, robots are gradually evolving from traditional closed automated equipment into embodied intelligent agents capable of collaborating with humans in shared physical spaces. These robots need to continuously perceive environmental conditions, understand human intentions, and make safety control decisions during task execution. Their safety control capabilities have become a key technological foundation for the widespread application of human-robot collaborative systems. Industrial robots, as a typical application of embodied intelligent agents in actual production environments, face even higher demands for the real-time performance, intelligence, and reliability of safety control in open collaborative scenarios.

[0003] Robot safety control typically relies on pre-programmed logic or traditional visual monitoring, such as setting up fixed fences or using infrared sensors. However, these methods have poor anti-interference capabilities, struggle to understand behavioral intentions, and lack language understanding capabilities, requiring manual reprogramming for dynamic tasks. While multimodal large language models offer a new approach, general models generally lack the inherent modeling ability for physical constraints, safety rules, and real-time control requirements. Directly applying them to safety control scenarios can easily lead to decision-making biases and fail to meet the practical application requirements of low latency and high reliability. Therefore, there is an urgent need for a robot safety control method that incorporates multimodal perception, embodied cognition, and embedded safety constraint mechanisms. Summary of the Invention

[0004] This invention provides a robot safety control method based on a multimodal embodied intelligent agent. The technical problem it solves is that existing robot safety control methods have insufficient robustness of multimodal perception in complex environments, lack the ability to comprehensively reason about task instructions and environmental states, and have potential safety risks when using large-scale models for control decisions, which may lead to safety accidents or failure to efficiently complete tasks during robot operation.

[0005] To address the above technical problems, the first aspect of this invention provides a robot safety control method based on a multimodal embodied intelligent agent, comprising the following steps:

[0006] A robot safety control method based on a multimodal embodied intelligent agent includes the following steps:

[0007] S1: Collect environmental visual information, voice commands, and robot body state data through multimodal sensors on the robot, and perform spatiotemporal alignment and feature fusion on the original multimodal perception data; on this basis, construct a key entity encoding mechanism for robot operation scenarios, and map the fused multimodal features into key entity semantic representations that can be understood by the embodied intelligent agent.

[0008] S2: Based on the key entity semantic representation described in step S1, construct a multi-stage task reasoning process for high-level task instructions, parse the high-level task instructions and decompose them into several sub-task steps, introduce safety semantics based on human motion prediction in each stage or node of the task reasoning process, verify the executability of the current sub-task, and execute safety constraints based on the safety semantics, thereby generating control instructions that take into account both the operation objectives and safety requirements.

[0009] S3: The control commands described in step S2 are sent to the robot's actuators to perform actions, and a dynamic safety intervention mechanism based on environmental status, personnel status, and predicted risks is constructed. During the execution process, the dynamic changes in the environment and personnel status are perceived in real time, and the current execution action is evaluated in real time based on the prediction results of future status. When a potential safety risk is detected, the dynamic intervention logic is triggered, and the action execution results and status changes are fed back to step S1 to update the key entity semantic representation, thereby forming a closed-loop safety control of perception, reasoning, execution, and feedback.

[0010] A second aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the robot safety control method based on a multimodal embodied intelligent agent described above.

[0011] A third aspect of the present invention is to provide a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the robot safety control method based on a multimodal embodied intelligent agent described above.

[0012] The robot safety control method based on multimodal embodied intelligent agents provided by this invention constructs multimodal perception and semantic representation oriented towards key entities, enabling the robot to form a unified embodied internal state. On this basis, a multi-stage task reasoning mechanism for safety constraint injection is introduced, combined with prediction-driven dynamic safety execution and closed-loop feedback, to achieve safe, continuous and reliable control of embodied intelligent agents in complex human-machine collaborative environments. Attached Figure Description

[0013] Figure 1 This is a framework diagram of the robot safety control method based on multimodal embodied intelligent agents according to the present invention.

[0014] Figure 2 This is a flowchart of the task reasoning based on the thought chain in step S2 of the robot safety control method based on multimodal embodied intelligent agents of the present invention.

[0015] Figure 3 This is a time-series graph showing the changes in personnel displacement.

[0016] Figure 4 This is a diagram illustrating the process of safety risk assessment and threshold determination. Detailed Implementation

[0017] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0018] like Figure 1 As shown, this embodiment discloses a robot safety control method based on a multimodal embodied intelligent agent, including the following steps:

[0019] S1: Collect environmental visual information, voice commands, and robot body state data through multimodal sensors on the robot, and perform spatiotemporal alignment and feature fusion on the original multimodal perception data; on this basis, construct a key entity encoding mechanism for robot operation scenarios, and map the fused multimodal features into key entity semantic representations that can be understood by the embodied intelligent agent.

[0020] S2: Based on the key entity semantic representation described in step S1, construct a multi-stage task reasoning process for high-level task instructions, parse the high-level task instructions and decompose them into several sub-task steps, introduce safety semantics based on human motion prediction in each stage or node of the task reasoning process, verify the executability of the current sub-task, and execute safety constraints based on the safety semantics, thereby generating control instructions that take into account both the operation objectives and safety requirements.

[0021] S3: The control commands described in step S2 are sent to the robot's actuators to perform actions, and a dynamic safety intervention mechanism based on environmental status, personnel status, and predicted risks is constructed. During the execution process, the dynamic changes in the environment and personnel status are perceived in real time, and the current execution action is evaluated in real time based on the prediction results of future status. When a potential safety risk is detected, the dynamic intervention logic is triggered, and the action execution results and status changes are fed back to step S1 to update the key entity semantic representation, thereby forming a closed-loop safety control of perception, reasoning, execution, and feedback.

[0022] The following section will explain each step in detail.

[0023] Step S1 specifically includes the following steps:

[0024] S11. In this embodiment, the robot operation is perceived by the multimodal sensor. The multimodal sensor includes a vision sensor, a voice acquisition module and a body state sensor, which are used to collect visual information, voice commands and body state information such as robot pose, joint angle and motion state, respectively.

[0025] Suppose The raw multimodal sensing data collected at any time is represented as follows:

[0026] ,

[0027] in, Representing visual data, Indicates instruction data, This represents the robot's physical state data;

[0028] To improve the quality of the multimodal raw sensing data and the stability of subsequent processing, the multimodal raw sensing data is subjected to denoising, normalization and format unification preprocessing operations respectively.

[0029] In steps S12 and S11, the sampling frequencies and installation locations of different sensors differ. To eliminate the spatiotemporal inconsistencies in the multimodal raw sensing data, this embodiment performs time synchronization processing on the multimodal raw sensing data based on the timestamp information corresponding to each modality's raw sensing data, mapping the different modal raw sensing data to a unified time axis. For multimodal raw sensing data with inconsistent sampling frequencies, time synchronization of the multimodal raw sensing data is achieved through time window matching, interpolation, or resampling.

[0030] In the spatial dimension, spatial information is aligned based on sensor spatial calibration parameters, as shown in the following expression:

[0031] ,

[0032] in, This indicates the spatial position of a point in the sensor coordinate system. This indicates the position of a point in space within the robot's reference coordinate system. This represents the spatial transformation matrix from the sensor coordinate system to the robot reference coordinate system;

[0033] S13. Extract the corresponding modal features from the aligned visual data, voice command data, and robot body state data described in step S12, respectively. The process is as follows:

[0034] ,

[0035] ,

[0036] ,

[0037] in, , , These represent the feature extraction functions for the corresponding modes.

[0038] A cross-modal association modeling mechanism is used to fuse features from different modalities to establish semantic relationships between information from different modalities. The fusion result is represented as follows:

[0039] ,

[0040] in, This represents the fused multimodal features;

[0041] S14. Based on the fused multimodal features described in step S13, identify key entities and human-computer interaction objects in the robot's operating environment:

[0042] Firstly, based on visual features The system detects or segments personnel, robots, or tools in the environment to obtain a set of candidate key entities. Based on this, voice command features are introduced. and robot body state characteristics Cross-modal semantic matching and contextual analysis are performed to eliminate ambiguity or occlusion problems that may arise from single visual perception. The category attributes and physical interaction semantics of the candidate key entities are supplemented with annotations to determine the attribute labels of each key entity. ;

[0043] Subsequently, using the visual detection results of multiple consecutive frames, the spatial position and motion state of each key entity are tracked and evaluated in the robot reference coordinate system, thereby obtaining the spatial position of each key entity in the robot reference coordinate system. and motion state information of key entities ;

[0044] Furthermore, for personnel entities, their skeleton or action feature sequences within a continuous time window are extracted, and their current action or operation semantics are extracted through spatiotemporal correlation modeling. This reflects the dynamic behavioral trends of workers in the current environment;

[0045] Specifically, for a human entity, the skeletal or action feature sequence within a continuous time window is extracted. Spatially, based on the physical connection topology of the human skeleton, the spatial dependencies between joint nodes are calculated to capture the local structural features of human posture. Temporally, the motion trajectory and dynamic change rate of skeletal joints between adjacent time frames are calculated to capture the global evolution trend of limb movements. Subsequently, the spatial structural features and temporal evolution features are fused and decoded to map the continuous high-dimensional skeletal trajectory sequence into low-dimensional discrete semantic labels, thereby generating action or operational semantics that reflect the dynamic behavioral trends of the worker in the current environment. .

[0046] The attribute information, spatial location information, and interaction relationships between the key entities and robots or personnel are encoded.

[0047] For the The key entity state vector of each identified key entity is defined as follows:

[0048] ,

[0049] in, This indicates the spatial position of the key entity in the robot's reference coordinate system. This represents the motion state information of key entities. This indicates the corresponding action or operation semantics. Attribute tags representing key entities, used to distinguish between people, robots, or tools;

[0050] S15. Based on the encoding results of the key entities in step S14, construct a set of key entity semantic representations for robot control:

[0051] ,

[0052] in, This indicates the number of key entities identified in the current environment. The key entity semantic representation is used to describe the environmental state, object relationships, and operational semantics related to the current task, and serves as input for subsequent task reasoning and security control.

[0053] In this embodiment, step S2 is performed based on the key entity semantic representation constructed in step S1, such as... Figure 2 As shown, the chained reasoning and security control decision-making for high-level task instructions includes the following steps:

[0054] S21. Receive high-level task instructions. Combined with the key entity semantic representation set described in step S15 The high-level task instructions are parsed to determine the task intent and the operation objects and constraints related to the current environmental state.

[0055] Specifically, the high-level task instructions are first semantically decomposed and transformed into structured semantic units, including action semantics, target object semantics and constraint semantics. The action semantics are used to describe the task execution type, the target object semantics are used to indicate the object on which the task acts, and the constraint semantics are used to limit the task execution conditions.

[0056] Based on this, the semantics of the target object obtained from the analysis will be combined with the set of key entity semantic representations. Attribute tags of key entities in the middle Matching is performed to identify key target entities relevant to the task. Simultaneously, the spatial location of each key entity is considered. and motion status information A correlation analysis was conducted to examine the feasibility of the task and the environmental constraints.

[0057] Specifically, it involves combining the parsed target object semantics with the entity semantic representation set. Attribute tags of each entity Perform matching; simultaneously, consider the spatial location of each entity. and motion status information The matched entities are subjected to out-of-control state verification. Based on the verification results, target entities that meet the requirements of operational reachability and state stability are selected to form a target object set. Furthermore, the physical constraints implicit in the spatial position and motion state information are extracted and added to the constraint set. middle.

[0058] The above parsing results can be represented as a structured task semantic vector:

[0059] ,

[0060] in, Indicates action semantics, This represents the set of target objects associated with the current task. The structured task semantic vector represents a set of constraints and is used to uniformly represent the execution requirements of the current task at the semantic level, and serves as the input for the subsequent thought chain reasoning process.

[0061] S22. Based on the stated task intent, the high-level task is broken down into several sub-tasks with a sequential order, and the sub-tasks are organized into an ordered thought chain:

[0062] ,

[0063] Each node in the thought chain The state of a subtask is represented by:

[0064] ,

[0065] in, Subtasks The semantic subset of the key entities associated with the relationship This represents the goal constraint of a subtask, used to describe the operational objectives of that stage;

[0066] In this embodiment, the thought chain unfolds step by step with sub-tasks as nodes, and the reasoning result of the previous node serves as the input of the next node, thereby forming a progressive and reversible chain reasoning structure.

[0067] S23. For personnel entities involved in the thought chain, based on their current key entity semantics, predict their spatial state at future moments to assess potential security risks.

[0068] For personnel entities Its future The predicted position at that time is:

[0069] ,

[0070] in, Indicates the current location of the personnel. Indicates the current trend of personnel movement. Representation based on personnel action semantics And the action offset obtained from its temporal evolution prediction.

[0071] By using the above methods, the results of human motion prediction are mapped onto the future spatial state, providing a spatiotemporal basis for safety risk assessment.

[0072] S24. At the current stage of the task reasoning process, introduce safety semantics based on human action prediction to verify the executability of the current sub-task. When potential security risks are detected, filter, adjust or replace the sub-task steps or reasoning path.

[0073] Targeting the nodes of the thinking chain The corresponding security risk assessment value is defined as follows:

[0074] ,

[0075] in, Represents a set of personnel entities; This represents the set of key parts of the robot; Indicates the critical parts of the robot at any time Spatial location; The risk weights between personnel and robot parts are represented by personnel action semantics. Decide; This represents the distance-risk mapping function, used to map spatial distance to risk values; Indicates Euclidean distance;

[0076] In the process of reasoning through thought chains, for each node Perform the following safety feasibility assessment:

[0077] ,

[0078] in, It is a preset security risk threshold;

[0079] When the judgment result does not meet the above conditions, the current thought chain node is considered infeasible, and the system rolls back, adjusts or regenerates the current subtask, thereby proactively avoiding potential security risks during the reasoning phase.

[0080] S25. When a certain thought chain node When safety constraints are met, the sub-task objective corresponding to that node is used. Based on the semantic state of key entities, control instructions are generated to guide the robot's execution, and the control instructions are output.

[0081] Combination Figure 3 and Figure 4 The effectiveness of the closed-loop mechanism for spatiotemporal state prediction and security risk assessment in step S2 above will be further explained. For example... Figure 3 As shown, the spatial prediction trajectory of workers based on human action semantics exhibits a smooth evolution trend without abnormal fluctuations, effectively suppressing state transitions and thus providing a highly reliable spatiotemporal prior basis for subsequent risk calculation; furthermore, as Figure 4 As shown, based on the accurate predictive input, the risk assessment value can respond in real time to the dynamic changes in the human-machine collaboration state. When the risk value exceeds the preset safety threshold, the system precisely triggers the backtracking and correction of the thought chain reasoning path, enabling the risk value to quickly fall back and remain below the safety threshold. The synergistic effect of these two mechanisms enables the dynamic identification and early blocking of potentially unsafe behaviors, significantly improving the embodied intelligent agent's proactive safety control capabilities in complex human-machine collaboration environments.

[0082] In this embodiment, step S3 receives the control decision and sub-task objective output from step S2. Without re-performing task reasoning, it achieves safe execution of the embodied intelligent agent at the physical level through risk cost modeling, ergonomic evaluation, and predictive optimization control. Specifically, it includes the following steps:

[0083] S3. Before the robot executes control commands, based on the key entity semantics and safety semantics obtained in steps S1 and S2, construct the execution layer execution cost function for the current execution action;

[0084] For any candidate motion edge in the execution trajectory Its execution cost is defined as:

[0085] ,

[0086] in, Based on execution cost, used to describe execution efficiency, it is specifically defined as:

[0087] ,

[0088] in, Indicates displacement distance. Indicates the execution time. Indicating the energy consumption during execution (such as the torque or current integral of a joint motor), in the two formulas above, , , For the corresponding weighting coefficients, This is a safety risk item, consisting of ergonomic risks and spatial risks, specifically defined as:

[0089] ,

[0090] in, This document describes a continuous risk score based on OWAS (Ovako Working Posture Analysis System) derived from temporal information analysis of human key points, corresponding to human posture, load, and motion semantics. OWAS is an ergonomic assessment method that classifies human posture and stress state and assigns risk levels to characterize human load and potential injury risks during work. In this embodiment, the OWAS method is extended to enable continuous risk scoring based on temporal data of human key points. This makes it more suitable for security assessment in dynamic human-machine collaboration scenarios;

[0091] and These represent the spatial positions of the personnel and the robot at the predicted time, respectively. The predicted personnel locations from step S23 can be directly referenced. ; This is a non-linear mapping function from distance to collision risk, typically set as an exponential or inverse proportional function where the function value increases as the distance decreases. and This is a weighting coefficient used to dynamically balance efficiency and safety in different work scenarios.

[0092] S32. During execution, the robot does not directly execute control commands at a single instant, but rather achieves predictive control by solving for the optimal control sequence within a finite time domain based on the current state and execution cost.

[0093] Construct the following optimal control model:

[0094] ,

[0095] in, Indicates the current execution state of the robot; Indicates the first The control input sequence in the step; The execution cost in step S31; This is the terminal safety cost function, used to constrain the robot's state at the end of the predicted time domain, preventing it from entering high-risk areas or unrecoverable singular states.

[0096] When the security risks of displaying future states are mentioned Exceeding the preset security threshold At that time, the implementation level adopts a tiered intervention strategy based on the risk level. Specific strategies include:

[0097] 1) Low-risk level: Smoothly adjust the execution speed or acceleration, for example, by reducing motion gain to increase operational margin;

[0098] 2) Medium risk level: Switch to human-robot collaboration mode, the robot changes from a dominant role to an auxiliary role, and performs impedance control by following the movement of the personnel;

[0099] 3) High risk level: Pause the current action and lock it, waiting for personnel status to change or the risk to be eliminated. If the risk persists, trigger an emergency stop or mission rollback.

[0100] After the action is completed, the actual execution trajectory, the real control input, and the monitored changes in ergonomic risks are fed back to step S1. These measured data are used to update the key entity semantic representations in step S1, particularly correcting parameter errors in the human motion model (such as...). This closed-loop feedback mechanism enables the embodied agent to continuously calibrate its understanding of the environmental state and human interaction habits as the task progresses, providing more accurate input for subsequent task reasoning and control decisions.

[0101] In summary, the present invention provides a robot safety control method based on a multimodal embodied intelligent agent. By constructing multimodal perception and semantic representation oriented towards key entities, the robot forms a unified embodied internal state. On this basis, a multi-stage task reasoning mechanism for safety constraint injection is introduced. Combined with prediction-driven dynamic safety execution and closed-loop feedback, the embodied intelligent agent achieves safe, continuous and reliable control in complex human-machine collaborative environments.

[0102] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for robot safety control based on multi-modal embodied agent, characterized in that, Includes the following steps: S1: Collect environmental visual information, voice commands, and robot body state data through multimodal sensors on the robot, and perform spatiotemporal alignment and feature fusion on the original multimodal perception data; on this basis, construct a key entity encoding mechanism for robot operation scenarios, and map the fused multimodal features into key entity semantic representations that can be understood by the embodied intelligent agent. S2: Based on the key entity semantic representation described in step S1, construct a multi-stage task reasoning process for high-level task instructions, parse the high-level task instructions and decompose them into several sub-task steps, introduce safety semantics based on human motion prediction in each stage or node of the task reasoning process, verify the executability of the current sub-task, and execute safety constraints based on the safety semantics, thereby generating control instructions that take into account both the operation objectives and safety requirements. S3: The control commands described in step S2 are sent to the robot's actuators to perform actions, and a dynamic safety intervention mechanism based on environmental status, personnel status, and predicted risks is constructed. During the execution process, the dynamic changes in the environment and personnel status are perceived in real time, and the current execution action is evaluated in real time based on the prediction results of future status. When a potential safety risk is detected, the dynamic intervention logic is triggered, and the action execution results and status changes are fed back to step S1 to update the key entity semantic representation, thereby forming a closed-loop safety control of perception, reasoning, execution, and feedback. 2.The multi-modal embodied-agent based robot safety control method of claim 1, wherein, Step S1 specifically includes the following steps: S11, provided in The multi-modal raw perception data collected at the moment is represented as: , wherein, represents visual data, represents voice instruction data, represents robot body state data; a preprocessing operation of denoising, normalizing and format unifying on the multi-modal raw perception data respectively; S12. Based on the timestamp information corresponding to the original sensing data of each modality, perform time synchronization processing on the original sensing data of the multimodality and map the original sensing data of different modalities to a unified time axis; for the original sensing data of the multimodality with inconsistent sampling frequencies, time synchronization of the original sensing data of the multimodality is achieved through time window matching, interpolation or resampling. In the spatial dimension, spatial information is aligned based on sensor spatial calibration parameters, as shown in the following expression: , wherein, represents a spatial point position in the sensor coordinate system, represents a spatial point position in the robot reference coordinate system, represents a spatial transformation matrix from the sensor coordinate system to the robot reference coordinate system; S13. Extract the corresponding modal features from the aligned visual data, voice command data, and robot body state data described in step S12, respectively. The process is as follows: , , , in, , , These represent the feature extraction functions for the corresponding modes; A cross-modal association modeling mechanism is used to fuse features from different modalities to establish semantic relationships between information from different modalities. The fusion result is represented as follows: , in, This represents the fused multimodal features; S14. Based on the fused multimodal features described in step S13, identify key entities and human-machine interaction objects in the robot's working environment, and encode the attribute information, spatial location information, and interaction relationship between the key entities and the robot or personnel. For the The key entity state vector of each identified key entity is defined as follows: , in, This indicates the spatial position of the key entity in the robot's reference coordinate system. This represents the motion state information of key entities. This indicates the corresponding action or operation semantics. Attribute tags representing key entities, used to distinguish between people, robots, or tools; S15. Based on the encoding results of the key entities in step S14, construct a set of key entity semantic representations for robot control: , in, This indicates the number of key entities identified in the current environment. The key entity semantic representation is used to describe the environmental state, object relationships, and operational semantics related to the current task, and serves as input for subsequent task reasoning and security control.

3. The robot safety control method based on a multimodal embodied intelligent agent as described in claim 1, characterized in that, Step S2 specifically includes the following steps: S21. Receive high-level task instructions, parse the high-level task instructions, and determine the operation objects and constraints related to the task intent and the current environmental state. S22. Based on the stated task intent, the high-level task is broken down into several sub-tasks with a sequential order, and the sub-tasks are organized into an ordered thought chain: , Each node in the thought chain The state of a subtask is represented by: , in, Subtasks The semantic subset of the key entities associated with the relationship This represents the goal constraint of a subtask, used to describe the operational objectives of that stage; S23. For personnel entities involved in the thought chain, based on their current key entity semantics, predict their spatial state at future moments to assess potential security risks. For personnel entities Its future The predicted position at that time is: , in, Indicates the current location of the personnel. Indicates the current trend of personnel movement. Representation based on personnel action semantics And the action offset obtained from its temporal evolution prediction.

4. The robot safety control method based on a multimodal embodied intelligent agent as described in claim 3, characterized in that, Step S2 further includes the following steps: S24. At the current stage of the task reasoning process, introduce safety semantics based on human action prediction to verify the executability of the current sub-task. When potential security risks are detected, filter, adjust or replace the sub-task steps or reasoning path. Targeting the nodes of the thinking chain The corresponding security risk assessment value is defined as follows: , in, Represents a set of personnel entities; This represents the set of key parts of the robot; Indicates the critical parts of the robot at any time Spatial location; The risk weights between personnel and robot parts are represented by personnel action semantics. Decide; This represents the distance-risk mapping function, used to map spatial distance to risk values; Indicates Euclidean distance; In the process of reasoning through thought chains, for each node Perform the following safety feasibility assessment: , in, It is a preset security risk threshold; When the judgment result does not meet the above conditions, the current thought chain node is considered infeasible, and the current subtask is rolled back, adjusted or regenerated, thereby proactively avoiding potential security risks during the reasoning stage. S25. When a certain thought chain node When safety constraints are met, the sub-task objective corresponding to that node is used. Based on the semantic state of key entities, control instructions are generated to guide the robot's execution, and the control instructions are output.

5. The robot safety control method based on a multimodal embodied intelligent agent as described in claim 1, characterized in that, Step S3 specifically includes the following steps: S31. Before the robot executes control commands, construct an execution layer execution cost function for the currently executed action; For any candidate motion edge in the execution trajectory Its execution cost is defined as: , in, Based on execution cost, used to describe execution efficiency, it is specifically defined as: , in, Indicates displacement distance. Indicates the execution time. Indicating energy consumption, in the two formulas above, , , For the corresponding weighting coefficients, This is a safety risk item, consisting of ergonomic risks and spatial risks, specifically defined as: , in, The OWAS continuous risk score is obtained based on the analysis of human body key point temporal information, corresponding to human posture, load and action semantics. and These represent the spatial positions of the personnel and the robot at the predicted time, respectively. This is a nonlinear mapping function from distance to collision risk; and This is a weighting coefficient used to dynamically balance efficiency and safety in different work scenarios.

6. The robot safety control method based on a multimodal embodied intelligent agent as described in claim 5, characterized in that, Step S3 further includes the following steps: S32. During execution, the robot does not directly execute control commands at a single instant, but rather achieves predictive control by solving for the optimal control sequence within a finite time domain based on the current state and execution cost. Construct the following optimal control model: , in, Indicates the current execution state of the robot; Indicates the first The control input sequence in the step; The execution cost in step S31; This is the terminal safety cost function, used to constrain the robot's state at the end of the predicted time domain, preventing it from entering high-risk areas or unrecoverable singular states.

7. The robot safety control method based on a multimodal embodied intelligent agent as described in claim 6, characterized in that, Step S32 further includes: When displaying security risks in future states Exceeding the preset security threshold At that time, the implementation level adopts a tiered intervention strategy based on the risk level. Specific strategies include: 1) Low-risk level: Smoothly adjust the execution speed or acceleration; 2) Medium risk level: Switch to human-robot collaboration mode, the robot changes from a dominant role to an auxiliary role, and performs impedance control by following the movement of the personnel; 3) High risk level: Pause the current action and lock it, waiting for personnel status to change or the risk to be eliminated. If the risk persists, trigger an emergency stop or mission rollback.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the robot safety control method based on a multimodal embodied intelligent agent as described in any one of claims 1 to 7.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the robot safety control method as described in any one of claims 1 to 7.