Embodied-agent-based risk avoidance method and device
Patent Information
- Application Number
- CN202611330021.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本申请的主要目的在于提供一种基于具身智能体的风险规避方法及装置,可以解决现有技术中动态环境风险预见性与实时规避能力不足的技术问题
[0017]本申请提供了一种基于具身智能体的风险规避方法,对环境感知数据进行时空配准与特征提取,生成环境状态信息;结合任务目标规划出初始动作序列;在执行每个动作单元前,联合环境与本体状态预测交互结果,并计算风险评估指标;若风险超出容错阈值,则实时修正原动作序列,生成并执行安全的替代动作序列。通过上述方式,本方法实现了在动态环境中对潜在风险的超前感知与主动规避,增强了具身智能体任务执行的鲁棒性与安全性。
Smart Images

Figure CN122838779A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of embodied intelligent agents, and in particular to risk avoidance methods and apparatus based on embodied intelligent agents. Background Technology
[0002] With the widespread application of embodied intelligent agents in complex scenarios such as warehousing and logistics, service navigation, and autonomous driving, higher demands are placed on their adaptability to dynamic environments and operational safety. Existing technologies typically rely on pre-defined path planning or simple reactive obstacle avoidance. These methods are often based on static environment assumptions or delayed sensor feedback, making it difficult to effectively predict and assess potential risks such as sudden obstacles, uncertain movements of interactive objects, and abnormal agent states. Summary of the Invention
[0003] The main objective of this application is to provide a risk avoidance method and device based on embodied intelligent agents, which can solve the technical problems of insufficient predictability and real-time avoidance capability of dynamic environmental risks in the prior art.
[0004] To achieve the above objectives, this application provides a risk avoidance method based on embodied intelligent agents, the method comprising: Spatiotemporal registration and feature extraction are performed on the environmental perception data of the embodied intelligent agent to obtain environmental state information; Based on the environmental state information and preset task target data, the current task sequence of the embodied intelligent agent is planned to obtain the planned action sequence; Before executing each action unit in the planned action sequence, the interaction state between the agent and the environment after executing the current action unit is predicted by combining the environmental state information with the ontological state data of the embodied agent, and a risk assessment index is calculated based on the interaction state and the ontological state data. When the risk assessment index exceeds the preset fault tolerance threshold, the planned action sequence is corrected in real time to generate an alternative action sequence. The embodied agent is controlled to avoid risks based on the alternative action sequence.
[0005] In one embodiment, the step of performing spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information includes: Spatiotemporal registration is performed on the raw environmental perception data collected by sensors deployed in multiple locations on the embodied intelligent agent to obtain a fused environmental perception data stream; Extract the dynamic obstacle features from the fused environmental perception data stream, and obtain an obstacle motion trajectory prediction set based on the dynamic obstacle features; A real-time environment map is constructed based on the fused environment perception data stream and the obstacle motion trajectory prediction set. Environmental status information is generated based on the obstacle motion trajectory prediction set and the real-time environment map.
[0006] In one embodiment, the step of planning the current task sequence of the embodied intelligent agent based on the environmental state information and preset task target data to obtain a planned action sequence includes: Decompose the preset task target data into ordered subtask nodes; The unexecuted subtask nodes are determined based on the ordered subtask nodes and the current task sequence of the embodied agent; Based on the environmental state information, a search is performed on the reachable paths corresponding to the unexecuted sub-task nodes to obtain a set of candidate paths; The candidate path set is filtered to determine the execution path; The execution path is transformed into a planned sequence of actions.
[0007] In one embodiment, the steps of predicting the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the agent's ontological state data, and calculating a risk assessment index based on the interaction state and the ontological state data, include: Obtain the ontological state data of the embodied intelligent agent; Predict the pose data of the embodied agent after the current action unit is executed; The pose data is compared with the obstacle information in the environmental state information to obtain the position interaction information; The pose data is compared with the obstacle motion trajectory prediction set to obtain trajectory interaction information; The interaction state between the embodied intelligent agent and the environment is obtained based on the location interaction information and the trajectory interaction information; Risk assessment indicators are calculated based on the interaction state and the ontology state data.
[0008] In one embodiment, the step of calculating the risk assessment index based on the interaction state and the ontology state data includes: The difference between the actual parameter values in the ontology state data and the preset threshold corresponding to the expected state is analyzed to obtain the ontology anomaly coefficient. The predicted distance and relative speed between the embodied intelligent agent and the nearest obstacle are determined based on the interaction state. The collision risk coefficient is calculated based on the deviation between the predicted distance and the safe distance, and the relative speed. The collision risk coefficient and the body anomaly coefficient are weighted and fused to obtain a risk assessment index.
[0009] In one embodiment, the step of real-time correction of the planned action sequence and generation of an alternative action sequence when the risk assessment index exceeds a preset fault tolerance threshold includes: When the risk assessment indicator exceeds a preset fault tolerance threshold, the risk type and degree of exceeding the limit of the risk assessment indicator are determined. Based on the risk type and the degree of exceeding the limit, a correction strategy is obtained by matching from a preset strategy library; Identify the risky action units in the planned action sequence that generate the current risk; Based on the aforementioned correction strategy, the parameters of the risk action unit are corrected to obtain a corrected action sequence; The corrected action sequence is subjected to a coherence check. If the check passes, an alternative action sequence is generated based on the corrected action sequence.
[0010] In one embodiment, the risk type includes collision risk, body failure risk, or blockage risk, and the correction strategy includes deceleration strategy, stopping strategy, and detour strategy. The step of matching and obtaining the correction strategy from a preset strategy library based on the risk type and the degree of exceeding the limit includes: When the risk type is collision risk and the degree of exceeding the limit is in the first level range, a deceleration strategy is matched from the preset strategy library; When the risk type is collision risk and the degree of exceeding the limit is in the second level range, or when the risk type is body failure risk, a stop strategy is matched from the preset strategy library. When the risk type is path blocking risk and the degree of exceeding the limit is in the third level range, a detour strategy is matched from the preset strategy library, wherein the first level range is smaller than the second level range, and the second level range is smaller than the third level range.
[0011] In one embodiment, after the step of controlling the embodied agent to avoid risks based on the alternative action sequence, the method further includes: During the execution of the alternative action sequence or the planned action sequence, environmental status information, risk assessment indicators, actual action sequence and execution result data are continuously recorded to generate historical execution logs. The historical execution logs are analyzed at preset intervals to identify target patterns that trigger risk avoidance. Determine a task planning strategy for planning the current task sequence of the embodied intelligent agent; The task planning strategy is optimized based on the target pattern.
[0012] In one embodiment, the step of controlling the embodied agent to avoid risks based on the alternative action sequence includes: Convert the alternative action sequence into motor control commands; The motor control commands are sent to each actuator of the embodied intelligent agent so that each actuator can perform the corresponding risk avoidance.
[0013] Furthermore, to achieve the above objectives, this application also proposes a risk avoidance device based on an embodied intelligent agent, which includes: The perception fusion module is used to perform spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information. The task planning module is used to plan the current task sequence of the embodied intelligent agent based on the environmental state information and preset task target data, so as to obtain the planned action sequence. The risk projection module is used to predict the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the ontological state data of the embodied agent before executing each action unit in the planned action sequence, and to calculate the risk assessment index based on the interaction state and the ontological state data. The sequence correction module is used to correct the planned action sequence in real time and generate an alternative action sequence when the risk assessment index is greater than the preset fault tolerance threshold. The behavior execution module is used to control the embodied agent to avoid risks based on the alternative action sequence.
[0014] Furthermore, to achieve the above objectives, this application also proposes a risk avoidance device based on embodied intelligent agents. The risk avoidance device based on embodied intelligent agents includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the risk avoidance method based on embodied intelligent agents as described above.
[0015] In addition, to achieve the above objectives, the present invention also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the risk avoidance method based on embodied intelligent agents as described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the risk avoidance method based on embodied intelligent agents as described above.
[0017] This application provides a risk avoidance method based on embodied intelligent agents. It performs spatiotemporal registration and feature extraction on environmental perception data to generate environmental state information; plans an initial action sequence based on the task objective; before executing each action unit, it predicts the interaction result by combining the environmental and ontology states and calculates a risk assessment index; if the risk exceeds the fault tolerance threshold, it corrects the original action sequence in real time and generates and executes a safe alternative action sequence. Through this approach, this method achieves proactive risk perception and avoidance in dynamic environments, enhancing the robustness and safety of embodied intelligent agent task execution. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating an embodiment of the risk avoidance method based on embodied intelligent agents in this application. Figure 2 This is a schematic diagram of the risk assessment process of multi-source information fusion of an embodied intelligent agent based on an embodiment of the risk avoidance method of embodied intelligent agents in this application. Figure 3 This is a schematic diagram of the risk assessment and strategy modification process of an embodiment of the risk avoidance method based on embodied intelligent agents in this application. Figure 4 This is a schematic diagram of the module structure of the risk avoidance device based on embodied intelligent agents according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the risk avoidance method based on embodied intelligent agents in the embodiments of this application.
[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] The main solution of this application embodiment is to perform spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information; Based on the environmental state information and preset task target data, the current task sequence of the embodied intelligent agent is planned to obtain the planned action sequence; Before executing each action unit in the planned action sequence, the interaction state between the agent and the environment after executing the current action unit is predicted by combining the environmental state information with the ontological state data of the embodied agent, and a risk assessment index is calculated based on the interaction state and the ontological state data. When the risk assessment index exceeds the preset fault tolerance threshold, the planned action sequence is corrected in real time to generate an alternative action sequence. The embodied agent is controlled to avoid risks based on the alternative action sequence.
[0025] With the widespread application of embodied intelligent agents in complex scenarios such as warehousing and logistics, service navigation, and autonomous driving, higher demands are placed on their adaptability to dynamic environments and operational safety. Existing technologies typically rely on pre-defined path planning or simple reactive obstacle avoidance. These methods are often based on static environment assumptions or delayed sensor feedback, making it difficult to effectively predict and assess potential risks such as sudden obstacles, uncertain movements of interactive objects, and abnormal agent states.
[0026] This application provides a solution that performs spatiotemporal registration and feature extraction on environmental perception data to generate environmental state information; plans an initial action sequence based on the task objective; before executing each action unit, it predicts the interaction result by combining the environment and ontology state, and calculates a risk assessment index; if the risk exceeds the fault tolerance threshold, it corrects the original action sequence in real time, generating and executing a safe alternative action sequence. Through this approach, this method achieves proactive risk perception and avoidance in dynamic environments, enhancing the robustness and safety of embodied intelligent agents in task execution.
[0027] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, a risk avoidance device based on an embodied intelligent agent, etc. This embodiment does not specifically limit it. The following uses a risk avoidance device based on an embodied intelligent agent as an example to describe this embodiment and the following embodiments.
[0028] This application provides a risk avoidance method based on embodied intelligent agents, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the risk avoidance method based on embodied intelligent agents in this application.
[0029] In this embodiment, the risk avoidance method based on embodied intelligent agents includes steps S10~S40: Step S10: Spatiotemporal registration and feature extraction are performed on the environmental perception data of the embodied intelligent agent to obtain environmental state information.
[0030] It should be noted that environmental perception data refers to the raw data about the external environment collected by various sensors carried by the embodied intelligent agent, such as lidar, cameras, and millimeter-wave radar. Environmental state information refers to the structured environmental description formed after processing this raw data, which has a unified spatiotemporal reference and contains key features.
[0031] Understandably, the process begins by unifying sensor data from different sources and at different times into a single coordinate system and time series through sensor calibration and data synchronization technologies, forming a fused environmental perception data stream. Next, computer vision and signal processing algorithms are used to detect and track dynamic obstacles in real time from this data stream. Kalman filtering or more advanced prediction models are then used to predict their short-term trajectories, creating a trajectory prediction set. Simultaneously, based on static obstacle detection and prior map information, a real-time environmental map containing both static and dynamic elements is constructed or updated. Finally, by integrating the predicted trajectory set of dynamic obstacles with the real-time environmental map, environmental state information is generated for subsequent decision-making and planning. This information not only describes the instantaneous static structure of the environment but, more importantly, includes predictive judgments about its dynamic evolution.
[0032] In one feasible implementation, the step of performing spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information includes: Spatiotemporal registration is performed on the raw environmental perception data collected by sensors deployed in multiple locations on the embodied intelligent agent to obtain a fused environmental perception data stream; Extract the dynamic obstacle features from the fused environmental perception data stream, and obtain an obstacle motion trajectory prediction set based on the dynamic obstacle features; A real-time environment map is constructed based on the fused environment perception data stream and the obstacle motion trajectory prediction set. Environmental status information is generated based on the obstacle motion trajectory prediction set and the real-time environment map.
[0033] It should be noted that raw environmental perception data refers to unprocessed raw sensor readings, fused environmental perception data stream refers to a multi-source sensor data sequence after spatiotemporal alignment, dynamic obstacle features refer to information on the geometric, motion, and other attributes of moving objects identified from sensor data, obstacle trajectory prediction set refers to a collection of multiple predictions of possible paths of dynamic obstacles in the future based on current and historical states, and real-time environmental map refers to a scene model that integrates static structure and dynamic obstacle prediction states.
[0034] Understandably, spatiotemporal registration is fundamental to building reliable environmental perception. First, sensor calibration determines the precise extrinsic parameters (position and attitude) of each sensor (such as LiDAR, camera, and millimeter-wave radar) relative to the intelligent agent, as well as their respective timestamp synchronization. For heterogeneous data, such as fusing image pixels with LiDAR point clouds, coordinate transformation is needed to project the point cloud onto the image plane, or vice versa. The core transformation formula is:
[0035] in, Represents a point in the world coordinate system. The transformation matrix (provided by the odometry or GPS / IMU) represents the transformation from the agent's body coordinate system to the world coordinate system. This represents a fixed transformation from the sensor (such as camera) coordinate system to the body coordinate system (obtained through calibration). This indicates the final coordinates of the point in the sensor coordinate system. Through this transformation chain, all sensor data is unified into a common, global world coordinate system, and by aligning timestamps, a coherent, multimodal fused sensing data stream is formed.
[0036] Subsequently, dynamic obstacle feature extraction and trajectory prediction are performed on the fused data stream. Typically, object detection algorithms (such as deep learning-based YOLO and PointPillars) are used to identify obstacles in the data, and multi-object tracking algorithms (such as SORT and DeepSORT) are used to associate the same obstacle in consecutive frames to form a tracking trajectory. Based on this trajectory history, its future motion is predicted. A commonly used and efficient model is the linear Kalman filter, whose prediction steps are formulated as follows:
[0037]
[0038] in, and Represent Estimates of the state vector (e.g., position, velocity) and covariance matrix at each moment. It is a state transition model that describes how states evolve over time. This is the process noise covariance, representing the uncertainty of the model. Through iterative prediction, the possible state distribution of each obstacle over several future time steps can be obtained. Considering the uncertainty of motion patterns, multi-model filtering (such as IMM) or scene rule-based sampling may be used to generate a prediction set of obstacle motion trajectories containing multiple possible paths. Each of the trajectories It is a sequence of states at a future moment.
[0039] Finally, a real-time environment map is constructed, and the final environmental state information is generated. This real-time environment map is not merely a traditional occupancy grid map or point cloud map, but a spatiotemporal semantic map incorporating dynamic prediction information. The system constructs a local map centered on the agent's current location. The underlying layer of this map contains occupancy information for static obstacles (ground, walls, etc., segmented from the point cloud). Based on this, the obstacle movement trajectories predicted in the previous step are then analyzed. This can be overlaid onto a map as spatio-temporal volumes or probability fields. For example, a risk probability that varies over time can be calculated for each grid cell in the map. This probability can be aggregated from the probabilities of trajectories in the trajectory prediction set crossing the grid. The final generated environmental state information... It is a structured data set, which can be formally represented as:
[0040] in, Represents static environment information. It is a trajectory prediction set. It is the future moment The dynamic risk field is characterized.
[0041] Step S20: Based on the environmental state information and preset task target data, plan the current task sequence of the embodied intelligent agent to obtain a planned action sequence.
[0042] It should be noted that the preset task target data refers to the description of the superior task that the agent needs to complete, such as "navigate from point A to point B" or "grab a specified object". The current task sequence refers to a series of sub-tasks that are initially decomposed to achieve the goal. The planned action sequence refers to the sequence of low-level instructions output by the planner that the agent controller can directly execute, such as speed and steering angle.
[0043] Understandably, a hierarchical task and motion planning framework is typically adopted. First, the task planner generates a coarse-grained, logically correct current task sequence based on the static semantic information in the preset goal and environmental state, using a search algorithm or a method that satisfies temporal logic, such as [progress to the intersection, turn left, proceed straight to the target point]. Then, the motion planner combines detailed environmental state information, including predicted trajectories of dynamic obstacles, to perform fine-grained path or trajectory optimization for each subtask under the premise of satisfying dynamic constraints. Model predictive control algorithms are typically used to solve an optimal control problem in a finite time domain online. While minimizing the deviation from the reference path and control energy consumption, the predicted trajectories of dynamic obstacles are used as time-varying constraints to avoid collisions. Finally, a safe and smooth planned action sequence is output in the short term.
[0044] In one feasible implementation, the step of planning the current task sequence of the embodied intelligent agent based on the environmental state information and preset task target data to obtain a planned action sequence includes: Decompose the preset task target data into ordered subtask nodes; The unexecuted subtask nodes are determined based on the ordered subtask nodes and the current task sequence of the embodied agent; Based on the environmental state information, a search is performed on the reachable paths corresponding to the unexecuted sub-task nodes to obtain a set of candidate paths; The candidate path set is filtered to determine the execution path; The execution path is transformed into a planned sequence of actions.
[0045] It should be noted that the preset task target data refers to the high-level description of the final goal that the agent needs to complete, such as "send outdoor objects to the indoor stacking area". The ordered subtask nodes refer to the sequence of steps with logical or spatiotemporal order that are decomposed to achieve the final goal. The current task sequence refers to the list of subtask nodes that the agent is executing or has been queued. The unexecuted subtask nodes refer to the steps in the current task sequence that have not yet started to be processed. The reachable path refers to the collision-free path that connects the agent's current position and the subtask target point under the constraints of environmental state information. The candidate path set refers to the set of a series of optional paths that meet basic safety conditions obtained through search. The execution path refers to the final path selected from the candidate path set for actual navigation.
[0046] Understandably, breaking down high-level task objectives into ordered subtask nodes is the first step in planning, a process that typically relies on a predefined task tree or knowledge graph. For example, the task "transfer an outdoor object to an indoor storage area" might be broken down into [go to the outdoor object area, lift the object, go to the corresponding indoor area]. After identifying the unexecuted subtask nodes, the planner needs to perform a path search for the next sub-objective (such as "go to the outdoor object area"). Environmental state information should be taken into account. Includes static maps and dynamic obstacle trajectory prediction set Pathfinding algorithms (such as A and D Lite) are primarily used on static maps. This process is performed to quickly find geometrically connected reachable paths. However, considering only static obstacles is insufficient, thus generating an initial set of paths. Further evaluation of its correlation with dynamic obstacle prediction trajectories is needed. The spatiotemporal conflict. One evaluation method is to calculate the agent's movement along each candidate path. Expected trajectory during driving And check it against all Does a spatiotemporal intersection exist, i.e., does a point in time exist? This causes the agent's position to overlap with that of the obstacle:
[0047] This collision detection method can filter out paths with a high risk of collision, forming a set of feasible candidate paths.
[0048] Next, the final execution path is determined from the set of feasible candidate paths and transformed into a specific sequence of planned actions. Path selection is a multi-objective optimization decision-making process, whose objective function is... Typically, path length, smoothness, safety, and the risk cost associated with dynamic obstacles are considered comprehensively. The objective function can be expressed as:
[0049] in, It is the road geometric length, It is the smoothness cost related to path curvature. Based on prediction set Calculate the path risk cost (e.g., the reciprocal of the minimum distance to a dynamic obstacle, or the integral of traversing a high-risk area). Weights , , Used to balance the importance of different factors. By evaluating And select the path with the lowest cost as the execution path. Finally, the path This is transformed into a sequence of planned actions, typically using trajectory optimization algorithms (such as path trackers based on Model Predictive Control, MPC). The MPC controller solves a finite-time optimal control problem based on the current state and the path... Generate a series of control commands (such as linear velocity) and angular velocity (i.e., planning action sequences) This enables the intelligent agent to smoothly and safely track the execution path.
[0050] Step S30: Before executing each action unit in the planned action sequence, the interaction state between the agent and the environment after executing the current action unit is predicted by combining the environmental state information and the ontology state data of the embodied agent, and a risk assessment index is calculated based on the interaction state and the ontology state data.
[0051] It should be noted that an action unit refers to a basic control instruction in a planned action sequence, such as an instruction for a specific speed and direction. Ontology state data refers to the real-time attitude, speed, position and other internal state information of the agent itself. Interaction state refers to the new spatiotemporal relationship formed between the agent's own state and the environmental state (especially dynamic obstacles) after the agent performs an action. Risk assessment index refers to the numerical value or level used to quantify the potential danger level of this action execution.
[0052] Understandably, a forward simulation prediction model is typically constructed. In each control cycle, before issuing an action command, the planner uses the current environmental state information (including predicted trajectories of dynamic obstacles) and the agent's own state (such as position and velocity) as initial conditions. By embedding a simplified agent kinematics or dynamics model, the planner quickly simulates the short-term state evolution trajectory of the agent after executing the action unit. This predicted trajectory is then spatiotemporally overlapped with obstacles in the environment (especially the predicted trajectory set of dynamic obstacles). If a collision or entry into a dangerous area is predicted, a quantitative risk assessment index is calculated based on parameters such as minimum distance, relative speed, and time to reach the collision point.
[0053] In one feasible implementation, the steps of predicting the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the ontological state data of the embodied agent, and calculating the risk assessment index based on the interaction state and the ontological state data, include: Obtain the ontological state data of the embodied intelligent agent; Predict the pose data of the embodied agent after the current action unit is executed; The pose data is compared with the obstacle information in the environmental state information to obtain the position interaction information; The pose data is compared with the obstacle motion trajectory prediction set to obtain trajectory interaction information; The interaction state between the embodied intelligent agent and the environment is obtained based on the location interaction information and the trajectory interaction information; Risk assessment indicators are calculated based on the interaction state and the ontology state data. The step of calculating the risk assessment index based on the interaction state and the ontology state data includes: The difference between the actual parameter values in the ontology state data and the preset threshold corresponding to the expected state is analyzed to obtain the ontology anomaly coefficient. The predicted distance and relative speed between the embodied intelligent agent and the nearest obstacle are determined based on the interaction state. The collision risk coefficient is calculated based on the deviation between the predicted distance and the safe distance, and the relative speed. The collision risk coefficient and the body anomaly coefficient are weighted and fused to obtain a risk assessment index.
[0054] It should be noted that ontology state data refers to the real-time physical state parameter set of the agent itself; pose data refers to the predicted values of the agent's position and posture in space; obstacle information refers to the geometric contours and positions of static obstacles in the environment and the current state of dynamic obstacles; obstacle trajectory prediction set refers to the multimodal prediction results of the motion paths of dynamic obstacles in the environment over a period of time based on the current moment; position interaction information refers to the spatial relationship between the predicted pose of the agent and static obstacles; trajectory interaction information refers to the spatiotemporal relationship between the predicted trajectory of the agent and the predicted trajectory of dynamic obstacles; and interaction state is a description of the overall contact between the agent and the environment after integrating position interaction information and trajectory interaction information.
[0055] It should be understood that the actual parameter value refers to the value measured by the sensor in real time in the ontological state data; the expected state refers to the ideal state calculated based on the current action unit and control model; the preset threshold is the allowable deviation range set by the user; the difference analysis is the operation of calculating the degree of deviation between the actual value and the expected value; the ontological anomaly coefficient is a value that quantifies the degree of ontological state anomaly; the predicted distance is the estimated interval between the agent and the nearest obstacle when they are most likely to interact in the future; the relative speed is the velocity vector difference between the agent and the obstacle at the interaction point; the safe distance is the minimum interval preset to ensure safety; the deviation is the difference between the predicted distance and the safe distance; the collision risk coefficient is a value that quantifies the probability and severity of a collision; the weighted fusion is the calculation process of merging multiple coefficients into a comprehensive index according to specific weights; and the risk assessment index is the final output comprehensive risk assessment value used for decision-making.
[0056] Understandably, referring to Figure 2 Predicting the agent's pose data after executing the current action unit is a crucial first step. This is typically accomplished using a simplified kinematic or dynamic forward simulation model. For example, for a differential wheeled robot, given the current action unit (such as the left wheel speed...),... and right wheel speed and time step Kinematic models can be used to predict the pose at the next moment. The formula can be expressed as:
[0057] in, It is the pose in the current ontology state data. It's the wheelbase. It's angular velocity. It's linear velocity. This is the predicted pose. This will be compared with obstacle information in the environmental status information. Location interaction information is calculated... Minimum Euclidean distance to the profiles of all static obstacles To obtain, if If the distance is less than the obstacle's inflated layer distance, a static collision risk is considered to exist. Trajectory interaction information is obtained by transferring the agent's position from its current pose to... The predicted motion segments, and the dynamic obstacle trajectory prediction set Each trajectory in Perform spatiotemporal cross-detection to determine the time... Do they intersect?
[0058] After acquiring the interaction state, it is necessary to calculate the collision risk coefficient from the external environment and the ontology anomaly coefficient from the ontology's internal state anomalies in parallel. When calculating the collision risk coefficient, firstly, based on the trajectory interaction information, find the dynamic obstacle and its trajectory point with the highest probability of spatiotemporal intersection with the agent's predicted trajectory, and calculate the predicted distance at that potential conflict point. (Distance between the agent and the obstacle) and the magnitude of the relative velocity vector Safe distance It is usually a function related to speed, for example ,in It is the preset reaction time. This is the minimum buffer distance. Collision risk factor. It can be modeled as a function proportional to distance deviation and relative velocity, and can be expressed as:
[0059] in, and This is a weighting coefficient used to adjust the contribution of distance-related unsafety and relative speed to risk. When hour, A value of 0 indicates that there is no collision risk under the current prediction. The collision risk coefficient comprehensively reflects the probability and potential severity of a collision.
[0060] On the other hand, the calculation of the ontology anomaly coefficient focuses on the reliability of the agent's own state. Dissimilarity analysis compares actual parameter values in the ontology state data (such as the actually measured wheel speed). ) and the expected state to be achieved when executing the current action unit (such as instruction speed) This is done by... (for each key state parameter) The degree of difference It can be calculated as follows:
[0061] in, This is the preset threshold for the allowable deviation of this parameter. (Subject-specific anomaly system) It can then be defined as the weighted sum or maximum value of the differences among all key parameters: or Weight This indicates the degree to which different parameter anomalies affect the overall system stability. This coefficient reflects the internal inconsistencies of the system caused by actuator failure, sensor drift, or external disturbances.
[0062] Finally, the collision risk coefficient is calculated through weighted fusion. and ontological anomaly coefficient The results are combined into the final risk assessment indicator R. The fusion formula typically uses a linear weighted sum:
[0063] Weight and Calibration needs to be performed based on the specific application scenario; for example, in a congested dynamic environment, it may be assigned... Higher weighting, while more attention may be paid to tasks with high actuator precision requirements. Furthermore, to handle extreme situations, a nonlinear function can be introduced. For example, when any coefficient exceeds a certain emergency threshold, R is directly set to its maximum value. The resulting R is a unified metric that integrates external environmental threats and internal state anomalies, used by subsequent decision-making modules to determine whether to continue executing the current action, adjust the action, or trigger an emergency stop.
[0064] Step S40: When the risk assessment index is greater than the preset fault tolerance threshold, the planned action sequence is corrected in real time to generate an alternative action sequence.
[0065] It should be noted that the preset fault tolerance threshold refers to the pre-set threshold value used to judge whether the risk level is acceptable, real-time correction refers to dynamically adjusting the original plan based on the risk assessment results during the execution of the action, and alternative action sequence refers to a new and safe sequence of action instructions that is re-planned and generated to avoid risks.
[0066] Understandably, when the risk assessment index R exceeds the preset fault tolerance threshold T, it indicates that the current planned action sequence has unacceptable risks, triggering a real-time correction mechanism. Correction strategies typically include local replanning and global replanning. Local replanning immediately stops the execution of the current high-risk action unit and, within a shortened planning horizon (such as around the current position), uses algorithms such as Fast Random Tree (RRT) or Model Predictive Control (MPC) to re-search and generate a short safe trajectory and its corresponding alternative action sequence (such as deceleration, steering, or emergency braking), with the primary objective of avoiding identified risks (such as dynamic obstacles). If local adjustments cannot find a feasible solution, it may trigger a higher-level global task replanning, re-evaluating or adjusting the sequence of unexecuted sub-task nodes to ensure the task can continue safely.
[0067] In one feasible implementation, the step of real-time correction of the planned action sequence and generation of an alternative action sequence when the risk assessment index exceeds a preset fault tolerance threshold includes: When the risk assessment indicator exceeds a preset fault tolerance threshold, the risk type and degree of exceeding the limit of the risk assessment indicator are determined. Based on the risk type and the degree of exceeding the limit, a correction strategy is obtained by matching from a preset strategy library; Identify the risky action units in the planned action sequence that generate the current risk; Based on the aforementioned correction strategy, the parameters of the risk action unit are corrected to obtain a corrected action sequence; The corrected action sequence is checked for coherence. If the check passes, an alternative action sequence is generated based on the corrected action sequence. The risk types include collision risk, body failure risk, or blockage risk; the correction strategies include deceleration strategies, stopping strategies, and detour strategies; the step of matching and obtaining a correction strategy from a preset strategy library based on the risk type and the degree of exceeding the limit includes: When the risk type is collision risk and the degree of exceeding the limit is in the first level range, a deceleration strategy is matched from the preset strategy library; When the risk type is collision risk and the degree of exceeding the limit is in the second level range, or when the risk type is body failure risk, a stop strategy is matched from the preset strategy library. When the risk type is path blocking risk and the degree of exceeding the limit is in the third level range, a detour strategy is matched from the preset strategy library, wherein the first level range is smaller than the second level range, and the second level range is smaller than the third level range.
[0068] It should be noted that risk type refers to the main category of danger that causes the risk assessment index to exceed the limit; the degree of exceeding the limit refers to the amount and level by which the risk assessment index exceeds the preset fault tolerance threshold; the preset strategy library refers to the set of standard response plans for various risk situations that are stored in advance; risk action unit refers to the specific action instruction that is identified as directly causing the current high-risk state; correction strategy refers to the specific adjustment method to be taken to reduce risk; parameter correction refers to adjusting the values of control instructions such as speed, direction, and acceleration of the risk action unit; correction action sequence refers to the new action instruction sequence obtained after parameter adjustment; and continuity verification refers to the process of verifying whether the newly generated correction action sequence is smoothly connected with the previous and subsequent actions and meets the kinematic constraints of the agent.
[0069] Understandably, collision risk refers to the possibility of physical contact with an obstacle; agent failure risk refers to the possibility of performance degradation or loss of control due to hardware or software malfunctions; blockage risk refers to the possibility of the path forward being completely or partially blocked by static or dynamic obstacles; deceleration strategy refers to the strategy of extending reaction time and reducing impact energy by reducing movement speed; stopping strategy refers to the strategy of immediately suspending the current action and returning to a stationary state; detour strategy refers to the strategy of planning a new path to bypass the obstacle area; the first-level interval, the second-level interval, and the third-level interval refer to the risk severity level range divided according to the degree of exceeding the limit, and their numerical relationship represents the increase of risk.
[0070] In the specific implementation, refer to Figure 3 When the risk assessment index R exceeds the preset fault tolerance threshold T, the first step is to accurately diagnose the risk type (Type) and quantify the degree of exceedance (Severity). The risk type can be determined by analyzing the two main components that constitute the risk assessment index R—the collision risk coefficient. and ontological anomaly coefficient — Principal component analysis is used to determine this. For example, rules can be set: like (in If the threshold value is a proportional threshold (e.g., 0.7), then it is considered a collision risk. like If so, it is determined to be a risk of system failure; If the environmental status information indicates that the path ahead is occupied for a long time and local replanning has failed, it is judged as a risk of congestion.
[0071] The severity of exceeding the limit is usually graded based on the proportion of R exceeding T, and the calculation formula is as follows:
[0072] The results are then mapped to discrete level intervals, such as [0, 0.5T) for the first level interval (slight over-limit), [0.5T, 2T) for the second level interval (moderate over-limit), and [2T, ∞) for the third level interval (severe over-limit).
[0073] Next, based on the diagnosed Type and Severity, a correction strategy is matched from a pre-defined strategy library. The strategy library can be modeled as a lookup table or a rule engine. The matching process can be represented by a function:
[0074] And can be done according to the rules: when and , Output "Deceleration Strategy". This is because minor collision risks usually stem from an underestimated distance or an overestimated relative speed, and deceleration can effectively increase the safety margin.
[0075] when and ,or At any time (regardless of severity, the risk of system failure is high). The system outputs "Stop Policy". This indicates that the risk is high or the system itself is unreliable, and the primary task is to immediately terminate the dangerous action to prevent the situation from escalating.
[0076] when and , The output is "bypass strategy". This means that minor local adjustments (such as slowing down) are no longer sufficient to solve the problem, and a more fundamental path change is needed to bypass the congested area.
[0077] After determining the correction strategy, it is necessary to locate the specific risky action units in the planned action sequence that are causing the current high risk. This is typically achieved through sensitivity analysis or reverse tracing. One approach is to record each action unit during the prior risk assessment. Contribution to the overall risk indicator R A risk action unit can be defined as:
[0078] in This refers to the action to be executed in the sequence. Contribution. It can be calculated in the execution of the action unit The risk level can be estimated by measuring changes in risk indicators before and after the event, or by analyzing the correlation between parameters of the action unit (such as target speed) and risk factors (such as predicted distance). Precise positioning. This forms the basis for effective correction, ensuring that resources are focused on addressing the most critical sources of risk.
[0079] Then, based on the matched correction strategy, the risk action unit is... The parameters are corrected. Assuming the original action unit... Includes a set of control parameters (such as target speed) Target angular velocity Correction operation Modify these parameters according to the strategy type to generate a revised parameter set. For the "deceleration strategy", the correction function is:
[0080] in, It is a deceleration factor less than 1 (e.g., 0.5), where the speed is reduced while the direction remains unchanged. For the "stopping strategy," the correction function might be: This involves setting both velocity and angular velocity to 0. For the "detour strategy," the correction function is more complex; it might call a local path planner to generate a new short-term target point by bypassing the obstruction, starting from the current position, and then inversely derive a new velocity command from this. and ,Right now: .
[0081] After generating the corrected action sequence, a coherence check must be performed to ensure its feasibility. The coherence check mainly includes kinematic coherence and dynamic coherence. Kinematic coherence checks whether the corrected action sequence satisfies the agent's motion constraints; for example, for a differentially driven robot, whether consecutive poses satisfy nonholonomic constraints. A coherence loss function can be defined. :
[0082] in, It is the pose at time step t. It is a kinematic model. This is the corrected motion unit. Dynamic consistency checks whether acceleration and jerk are within allowable ranges to prevent abrupt changes in control commands that could lead to actuator saturation or instability. The condition for passing the check is... The value is less than the threshold and all dynamic parameters are within limits. If the verification fails, it is necessary to iteratively adjust the correction parameters (such as further reducing the rate of change of velocity) or backtrack to try other strategies until a feasible sequence is found.
[0083] Finally, after the coherence check passes, a final alternative action sequence is generated based on the revised action sequence. This process involves more than just replacing risky units in the original sequence; it may also require fine-tuning a small number of subsequent action units to ensure a smooth transition from the revised state back to the original planned path or mission objective. Alternative Action Sequence This can be formally expressed as:
[0084] in arrive It is an action unit that has been executed or is risk-free. It is a revised risk action unit. This is a follow-up action adjusted to maintain consistency. After generation, the alternative sequence is sent to the underlying controller for execution, while the monitoring system resets the risk assessment cycle and continuously tracks the execution effect of the alternative sequence, forming a closed-loop safety control process of "assessment-decision-correction-execution".
[0085] Step S50: Control the embodied intelligent agent to avoid risks based on the alternative action sequence.
[0086] It should be noted that risk avoidance refers to the ultimate control behavior of keeping the embodied agent away from or eliminating identified risks by executing alternative action sequences, thereby ensuring the agent's own safety and the continuation of the task.
[0087] Understandably, risk avoidance based on embodied intelligent agents using alternative action sequences hinges on translating safety decisions at the planning layer into precise physical control commands at the lower level. This is typically achieved through a hierarchical control system: a higher-level decision-making module sends validated alternative action sequences (such as target speed, steering angle, etc.) to the lower-level motion controllers (such as PID controllers or Model Predictive Controllers, MPCs). These controllers then calculate each action unit in the sequence into specific actuator control signals (such as motor torque, servo angle) in real time, driving the intelligent agent to perform corresponding avoidance actions such as deceleration, stopping, or steering. During execution, the system continuously monitors the agent's state and the environment's state, comparing it with the expected state of the alternative sequence. If new risks or execution deviations are detected, a new round of risk assessment and sequence correction may be triggered, forming a real-time perception-planning-action closed loop to ensure the effectiveness and adaptability of avoidance actions until the risk is eliminated or the mission objective is achieved.
[0088] In one feasible implementation, after the step of controlling the embodied agent to avoid risks based on the alternative action sequence, the method further includes: During the execution of the alternative action sequence or the planned action sequence, environmental status information, risk assessment indicators, actual action sequence and execution result data are continuously recorded to generate historical execution logs. The historical execution logs are analyzed at preset intervals to identify target patterns that trigger risk avoidance. Determine a task planning strategy for planning the current task sequence of the embodied intelligent agent; The task planning strategy is optimized based on the target pattern.
[0089] It should be noted that historical execution logs refer to structured data sets recorded in timestamp order, including environmental observation data, risk assessment results, actual control command sequences, and final state feedback. Target patterns refer to event sequences, combinations of environmental features, or decision logics identified from historical log data through statistical analysis or machine learning methods that lead to the recurrence of high-risk states. Task planning strategies refer to the rules, cost functions, or algorithmic models followed when generating specific action sequences from high-level task objectives. Optimization refers to adjusting the parameters, rules, or structure of the task planning strategy based on identified risk patterns, aiming to reduce the probability of similar risks occurring in future task executions from the outset.
[0090] In practical implementation, continuous recording is a crucial step in building the data foundation. During execution, the system captures and stores multimodal data at a fixed frequency (e.g., 10Hz), forming a historical execution log. Its data structure can be designed as a time-series database, where each data point includes a timestamp (ttt) and an environment state vector. (e.g., point cloud features from LiDAR, embedding vectors from camera images, dynamic obstacle lists), risk assessment metrics Planning or alternative action sequences and the actual sequence of actions performed and instant execution results (e.g., whether a collision occurred, whether the path was deviated from). The core challenge in generating logs lies in data alignment and compression. It's necessary to ensure that asynchronous data from different sensors and control loops are precisely synchronized in timestamps, and may employ sliding windows or event-triggered mechanisms to record critical segments to avoid data explosion. This detailed log provides rich material for subsequent offline analysis, enabling the system to review the complete context of each risk avoidance decision.
[0091] Periodic log analysis is central to the system's self-evolution. The system initiates an offline analysis process at preset intervals (e.g., every 100 tasks completed or every 24 hours). This process first preprocesses historical execution logs, such as data cleaning and feature extraction. Then, data analysis methods are used to identify "target patterns." For example, cluster analysis (such as DBSCAN) is used to identify patterns leading to high risk. ) environmental conditions The common characteristics of these are known as "risk environment patterns"; or specific action sequences can be discovered through association rule mining (such as the Apriori algorithm). This is followed by a "risk decision-making mode" where high-risk events frequently occur. Next, the system will review the current task planning strategy. This could be a rule-based decision tree or a reinforcement learning policy network. The optimization process involves adjusting the network based on the identified target pattern. For example, if a narrow corridor area is found to frequently trigger "collision risks," the penalty term for that area can be weighted in the cost function of the strategy; if a certain task startup order is found to easily lead to "blocking risks," the rule base of the strategy can be modified to prioritize other tasks. Through this closed loop of "practice-analysis-optimization," the task planning strategy can be continuously improved, shifting from passively responding to risks to actively avoiding risks, thereby enhancing the long-term autonomy and robustness of the agent.
[0092] In one feasible implementation, the step of controlling the embodied agent to avoid risks based on the alternative action sequence includes: Convert the alternative action sequence into motor control commands; The motor control commands are sent to each actuator of the embodied intelligent agent so that each actuator can perform the corresponding risk avoidance.
[0093] It should be noted that motor control commands refer to low-level electrical signals or digital commands that can be directly recognized and executed by the drivers of physical actuators such as motors and servos. They precisely specify the motion parameters of the actuators, such as speed, direction, angle, or stroke.
[0094] In practical implementation, the process of converting alternative action sequences into motor control commands relies on the motion control system within the embodied intelligent agent. This system first parses each abstract action in the sequence; for example, "avoiding to the left by 0.5 meters" is broken down into the target rotational speed and duration of the left and right wheels of the chassis. Then, the motion controller (such as a PID controller or model predictive controller) calculates the motor torque, voltage, or pulse width modulation signal required to accurately track the target trajectory based on the parsed target values and real-time states such as joint angles and wheel speeds fed back from sensors like encoders. These calculated low-level commands are sent in real-time to the drivers of each actuator via communication protocols such as CAN bus, Ethernet, or PWM interface. The drivers convert the received commands into high-voltage electrical signals, driving the motors or servos to precisely execute rotational or linear motion, thereby collaboratively completing the entire risk avoidance maneuver. The entire process requires extremely high real-time performance and reliability to ensure safety.
[0095] This embodiment provides a risk avoidance method based on embodied intelligent agents. It performs spatiotemporal registration and feature extraction on environmental perception data to generate environmental state information; plans an initial action sequence based on the task objective; before executing each action unit, it predicts the interaction results by combining the environmental and ontology states and calculates a risk assessment index; if the risk exceeds the fault tolerance threshold, it corrects the original action sequence in real time and generates and executes a safe alternative action sequence. Through this approach, the method achieves proactive risk perception and avoidance in dynamic environments, enhancing the robustness and safety of embodied intelligent agent task execution.
[0096] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the risk avoidance method based on embodied intelligent agents in this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0097] This application also provides a risk avoidance device based on an embodied intelligent agent; please refer to... Figure 4 Risk avoidance devices based on embodied intelligent agents include: The perception fusion module 10 is used to perform spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information. The task planning module 20 is used to plan the current task sequence of the embodied intelligent agent based on the environmental state information and preset task target data, so as to obtain a planned action sequence. The risk simulation module 30 is used to predict the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the ontological state data of the embodied agent before executing each action unit in the planned action sequence, and to calculate the risk assessment index based on the interaction state and the ontological state data. The sequence correction module 40 is used to correct the planned action sequence in real time and generate an alternative action sequence when the risk assessment index is greater than the preset fault tolerance threshold. The behavior execution module 50 is used to control the embodied intelligent agent to avoid risks based on the alternative action sequence.
[0098] In one feasible implementation, the perception fusion module 10 is further used to perform spatiotemporal registration on the raw environmental perception data collected by sensors arranged in multiple locations on the embodied intelligent agent to obtain a fused environmental perception data stream. Extract the dynamic obstacle features from the fused environmental perception data stream, and obtain an obstacle motion trajectory prediction set based on the dynamic obstacle features; A real-time environment map is constructed based on the fused environment perception data stream and the obstacle motion trajectory prediction set. Environmental status information is generated based on the obstacle motion trajectory prediction set and the real-time environment map.
[0099] In one feasible implementation, the task planning module 20 is further configured to decompose the preset task target data into ordered sub-task nodes; The unexecuted subtask nodes are determined based on the ordered subtask nodes and the current task sequence of the embodied agent; Based on the environmental state information, a search is performed on the reachable paths corresponding to the unexecuted sub-task nodes to obtain a set of candidate paths; The candidate path set is filtered to determine the execution path; The execution path is transformed into a planned sequence of actions.
[0100] In one feasible implementation, the risk inference module 30 is further used to acquire the ontological state data of the embodied intelligent agent; Predict the pose data of the embodied agent after the current action unit is executed; The pose data is compared with the obstacle information in the environmental state information to obtain the position interaction information; The pose data is compared with the obstacle motion trajectory prediction set to obtain trajectory interaction information; The interaction state between the embodied intelligent agent and the environment is obtained based on the location interaction information and the trajectory interaction information; Risk assessment indicators are calculated based on the interaction state and the ontology state data.
[0101] In one feasible implementation, the risk inference module 30 is further used to perform a difference analysis between the actual parameter values in the ontology state data and the preset threshold corresponding to the expected state to obtain the ontology anomaly coefficient. The predicted distance and relative speed between the embodied intelligent agent and the nearest obstacle are determined based on the interaction state. The collision risk coefficient is calculated based on the deviation between the predicted distance and the safe distance, and the relative speed. The collision risk coefficient and the body anomaly coefficient are weighted and fused to obtain a risk assessment index.
[0102] In one feasible implementation, the sequence correction module 40 is further configured to determine the risk type and degree of exceeding the limit of the risk assessment index when the risk assessment index is greater than a preset fault tolerance threshold; Based on the risk type and the degree of exceeding the limit, a correction strategy is obtained by matching from a preset strategy library; Identify the risky action units in the planned action sequence that generate the current risk; Based on the aforementioned correction strategy, the parameters of the risk action unit are corrected to obtain a corrected action sequence; The corrected action sequence is subjected to a coherence check. If the check passes, an alternative action sequence is generated based on the corrected action sequence.
[0103] In one feasible implementation, the sequence correction module 40 is further configured to match a deceleration strategy from a preset strategy library when the risk type is collision risk and the degree of exceeding the limit is in the first level range. When the risk type is collision risk and the degree of exceeding the limit is in the second level range, or when the risk type is body failure risk, a stop strategy is matched from the preset strategy library. When the risk type is path blocking risk and the degree of exceeding the limit is in the third level range, a detour strategy is matched from the preset strategy library, wherein the first level range is smaller than the second level range, and the second level range is smaller than the third level range.
[0104] In one feasible implementation, the behavior execution module 50 is further configured to continuously record environmental status information, risk assessment indicators, actual execution action sequences and execution result data during the execution of the alternative action sequence or planned action sequence, and generate historical execution logs. The historical execution logs are analyzed at preset intervals to identify target patterns that trigger risk avoidance. Determine a task planning strategy for planning the current task sequence of the embodied intelligent agent; The task planning strategy is optimized based on the target pattern.
[0105] In one feasible implementation, the behavior execution module 50 is further configured to convert the alternative action sequence into motor control commands; The motor control commands are sent to each actuator of the embodied intelligent agent so that each actuator can perform the corresponding risk avoidance.
[0106] The risk avoidance device based on embodied intelligent agents provided in this application, employing the risk avoidance method based on embodied intelligent agents in the above embodiments, can solve the technical problem of insufficient predictability and real-time avoidance capability of dynamic environment risks. Compared with the prior art, the beneficial effects of the risk avoidance device based on embodied intelligent agents provided in this application are the same as the beneficial effects of the risk avoidance method based on embodied intelligent agents provided in the above embodiments, and other technical features in the risk avoidance device based on embodied intelligent agents are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.
[0107] This application provides a risk avoidance device based on embodied intelligent agents. The risk avoidance device based on embodied intelligent agents includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the risk avoidance method based on embodied intelligent agents in the above embodiment 1.
[0108] The following is for reference. Figure 5 This document illustrates a structural schematic diagram of a risk avoidance device based on embodied intelligent agents suitable for implementing embodiments of this application. The risk avoidance device based on embodied intelligent agents in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The risk avoidance device based on embodied intelligent agents shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0109] like Figure 5As shown, the embodied agent-based risk avoidance device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the embodied agent-based risk avoidance device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the embodied agent-based risk avoidance device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows an embodied agent-based risk avoidance device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0110] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0111] The risk avoidance device based on embodied intelligent agents provided in this application, employing the risk avoidance method based on embodied intelligent agents in the above embodiments, can solve the technical problems of risk avoidance based on embodied intelligent agents. Compared with the prior art, the beneficial effects of the risk avoidance device based on embodied intelligent agents provided in this application are the same as the beneficial effects of the risk avoidance method based on embodied intelligent agents provided in the above embodiments, and other technical features in this risk avoidance device based on embodied intelligent agents are the same as the features disclosed in the method of the previous embodiment, and will not be repeated here.
[0112] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0113] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0114] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the risk avoidance method based on embodied intelligent agents in the above embodiments.
[0115] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0116] The aforementioned computer-readable storage medium may be included in a risk avoidance device based on an embodied intelligent agent; or it may exist independently and not be assembled into a risk avoidance device based on an embodied intelligent agent.
[0117] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by a risk avoidance device based on an embodied intelligent agent, cause the risk avoidance device based on an embodied intelligent agent to: perform spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information; Based on the environmental state information and preset task target data, the current task sequence of the embodied intelligent agent is planned to obtain the planned action sequence; Before executing each action unit in the planned action sequence, the interaction state between the agent and the environment after executing the current action unit is predicted by combining the environmental state information with the ontological state data of the embodied agent, and a risk assessment index is calculated based on the interaction state and the ontological state data. When the risk assessment index exceeds the preset fault tolerance threshold, the planned action sequence is corrected in real time to generate an alternative action sequence. The embodied agent is controlled to avoid risks based on the alternative action sequence.
[0118] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0120] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0121] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described risk avoidance method based on embodied intelligent agents, thereby solving the technical problem of risk avoidance based on embodied intelligent agents. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the risk avoidance method based on embodied intelligent agents provided in the above embodiments, and will not be repeated here.
[0122] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the risk avoidance method based on embodied intelligent agents as described above.
[0123] The computer program product provided in this application can solve the technical problem of insufficient predictability and real-time avoidance capability of dynamic environmental risks. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the risk avoidance method based on embodied intelligent agents provided in the above embodiments, and will not be repeated here.
[0124] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A risk avoidance method based on embodied intelligent agents, characterized in that, The risk avoidance method based on embodied intelligent agents includes: Spatiotemporal registration and feature extraction are performed on the environmental perception data of the embodied intelligent agent to obtain environmental state information; Based on the environmental state information and preset task target data, the current task sequence of the embodied intelligent agent is planned to obtain the planned action sequence; Before executing each action unit in the planned action sequence, the interaction state between the agent and the environment after executing the current action unit is predicted by combining the environmental state information with the ontological state data of the embodied agent, and a risk assessment index is calculated based on the interaction state and the ontological state data. When the risk assessment index exceeds the preset fault tolerance threshold, the planned action sequence is corrected in real time to generate an alternative action sequence. The embodied agent is controlled to avoid risks based on the alternative action sequence; The step of predicting the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the agent's ontological state data, and calculating the risk assessment index based on the interaction state and the ontological state data, includes: Obtain the ontological state data of the embodied intelligent agent; Predict the pose data of the embodied agent after the current action unit is executed; The pose data is compared with the obstacle information in the environmental state information to obtain the position interaction information; The pose data is compared with the obstacle motion trajectory prediction set to obtain trajectory interaction information; The interaction state between the embodied intelligent agent and the environment is obtained based on the location interaction information and the trajectory interaction information; Risk assessment indicators are calculated based on the interaction state and the ontology state data.
2. The method as described in claim 1, characterized in that, The steps of performing spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information include: Spatiotemporal registration is performed on the raw environmental perception data collected by sensors deployed in multiple locations on the embodied intelligent agent to obtain a fused environmental perception data stream; Extract the dynamic obstacle features from the fused environmental perception data stream, and obtain an obstacle motion trajectory prediction set based on the dynamic obstacle features; A real-time environment map is constructed based on the fused environment perception data stream and the obstacle motion trajectory prediction set. Environmental status information is generated based on the obstacle motion trajectory prediction set and the real-time environment map.
3. The method as described in claim 1, characterized in that, The step of planning the current task sequence of the embodied intelligent agent based on the environmental state information and preset task target data to obtain the planned action sequence includes: Decompose the preset task target data into ordered subtask nodes; The unexecuted subtask nodes are determined based on the ordered subtask nodes and the current task sequence of the embodied agent; Based on the environmental state information, a search is performed on the reachable paths corresponding to the unexecuted sub-task nodes to obtain a set of candidate paths; The candidate path set is filtered to determine the execution path; The execution path is transformed into a planned sequence of actions.
4. The method as described in claim 1, characterized in that, The step of calculating the risk assessment index based on the interaction state and the ontology state data includes: The difference between the actual parameter values in the ontology state data and the preset threshold corresponding to the expected state is analyzed to obtain the ontology anomaly coefficient. The predicted distance and relative speed between the embodied intelligent agent and the nearest obstacle are determined based on the interaction state. The collision risk coefficient is calculated based on the deviation between the predicted distance and the safe distance, and the relative speed. The collision risk coefficient and the body anomaly coefficient are weighted and fused to obtain a risk assessment index.
5. The method as described in claim 1, characterized in that, The step of real-time correction of the planned action sequence and generation of an alternative action sequence when the risk assessment index exceeds a preset fault tolerance threshold includes: When the risk assessment indicator exceeds a preset fault tolerance threshold, the risk type and degree of exceeding the limit of the risk assessment indicator are determined. Based on the risk type and the degree of exceeding the limit, a correction strategy is obtained by matching from a preset strategy library; Identify the risky action units in the planned action sequence that generate the current risk; Based on the aforementioned correction strategy, the parameters of the risk action unit are corrected to obtain a corrected action sequence; The corrected action sequence is subjected to a coherence check. If the check passes, an alternative action sequence is generated based on the corrected action sequence.
6. The method as described in claim 5, characterized in that, The risk types include collision risk, body failure risk, or blockage risk; the correction strategies include deceleration strategies, stopping strategies, and detour strategies; the step of matching and obtaining correction strategies from a preset strategy library based on the risk type and the degree of exceeding the limit includes: When the risk type is collision risk and the degree of exceeding the limit is in the first level range, a deceleration strategy is matched from the preset strategy library; When the risk type is collision risk and the degree of exceeding the limit is in the second level range, or when the risk type is body failure risk, a stop strategy is matched from the preset strategy library. When the risk type is path blocking risk and the degree of exceeding the limit is in the third level range, a detour strategy is matched from the preset strategy library, wherein the first level range is smaller than the second level range, and the second level range is smaller than the third level range.
7. The method as described in claim 1, characterized in that, Following the step of controlling the embodied agent to avoid risks based on the alternative action sequence, the method further includes: During the execution of the alternative action sequence or the planned action sequence, environmental status information, risk assessment indicators, actual action sequence and execution result data are continuously recorded to generate historical execution logs. The historical execution logs are analyzed at preset intervals to identify target patterns that trigger risk avoidance. Determine a task planning strategy for planning the current task sequence of the embodied intelligent agent; The task planning strategy is optimized based on the target pattern.
8. The method as described in claim 1, characterized in that, The steps of controlling the embodied agent to avoid risks based on the alternative action sequence include: Convert the alternative action sequence into motor control commands; The motor control commands are sent to each actuator of the embodied intelligent agent so that each actuator can perform the corresponding risk avoidance.
9. A risk avoidance device based on an embodied intelligent agent, characterized in that, The risk avoidance device based on embodied intelligent agents includes: The perception fusion module is used to perform spatiotemporal registration and feature extraction on the environmental perception data of the embodied intelligent agent to obtain environmental state information. The task planning module is used to plan the current task sequence of the embodied intelligent agent based on the environmental state information and preset task target data, so as to obtain the planned action sequence. The risk projection module is used to predict the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the ontological state data of the embodied agent before executing each action unit in the planned action sequence, and to calculate the risk assessment index based on the interaction state and the ontological state data. The sequence correction module is used to correct the planned action sequence in real time and generate an alternative action sequence when the risk assessment index is greater than the preset fault tolerance threshold. A behavior execution module is used to control the embodied agent to avoid risks based on the alternative action sequence; The step of predicting the interaction state between the agent and the environment after executing the current action unit by combining the environmental state information and the agent's ontological state data, and calculating the risk assessment index based on the interaction state and the ontological state data, includes: Obtain the ontological state data of the embodied intelligent agent; Predict the pose data of the embodied agent after the current action unit is executed; The pose data is compared with the obstacle information in the environmental state information to obtain the position interaction information; The pose data is compared with the obstacle motion trajectory prediction set to obtain trajectory interaction information; The interaction state between the embodied intelligent agent and the environment is obtained based on the location interaction information and the trajectory interaction information; Risk assessment indicators are calculated based on the interaction state and the ontology state data.