A body-aware intelligent safe-constrained action generation method and system based on world model forward reasoning
Patent Information
- Application Number
- CN202611255844.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明目的之一在于提供一种基于世界模型前瞻推理的具身智能安全约束动作生成方法,以解决现有技术中前瞻推演在原始观测空间进行而冗余度高,以及感知或通信失效时缺乏内生降级保护需依赖额外独立急停逻辑的问题
[0029]1、本发明将环境观测与本体状态融合为受物理约束的低维潜状态,并向其中注入当前构型下末端运动映射的局部敏感程度信息,使推演在剥离感知冗余的同时保留本体可操纵性,预测结果更贴合实际运动能力;同时通过在推演前将潜状态正交投影至由安全条件围成的合规区域,为整条推演链提供满足已建模安全要求的起点;且仅对越界状态作最小幅度修正,域内状态不予改动,避免将安全状态强制推至边界所造成的信息损失。
Smart Images

Figure CN122815917A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and robot control, and in particular to a method and system for generating embodied intelligent safety constraint actions based on world model forward reasoning. Background Technology
[0002] Embossed intelligent robots need to autonomously generate and execute actions in unstructured environments. To ensure the safety of these actions before execution, the industry commonly uses world models for forward-looking reasoning, predicting the future consequences of candidate actions and then deciding whether to allow them to proceed. However, existing technical solutions have limitations.
[0003] 1. Look-ahead simulations are often performed directly in the original observation space, resulting in high simulation overhead and potentially leading to predictions that exceed the local motion capabilities of the organism, lacking physical feasibility. 2. The initial state in look-ahead simulations lacks safety verification. Perception errors, feature mapping errors, or insufficient training data coverage can all cause the initial state itself to correspond to situations where the end effector intrudes into obstacles, joints exceed limits, or is too close to the workspace boundary. 3. Existing solutions mostly only output deterministic future states, failing to reflect the impact of model approximations and environmental unknowns on the prediction results.
[0004] Therefore, there is an urgent need for a security constraint algorithm system that can reduce the overhead of embodied intelligence inference and provide high accuracy. Summary of the Invention
[0005] One of the objectives of this invention is to provide a method for generating embodied intelligent safety constraint actions based on world model forward-looking reasoning, in order to solve the problems in the prior art where forward-looking inference is carried out in the original observation space with high redundancy, and where there is a lack of endogenous degradation protection when perception or communication fails, requiring the reliance on additional independent emergency stop logic.
[0006] This invention is achieved through the following technical solution: a method for generating embodied intelligent safety constraint actions based on world model forward reasoning, comprising the following steps: acquiring first observation data and second observation data, and determining a first state representation based on the first observation data and second observation data, wherein the first observation data is used to represent the self-motion state of the controlled device, the second observation data is used to represent the environment in which the controlled device is located, and the first state representation is a fusion representation of the first observation data and the second observation data in the latent space; based on the first state representation and the first action sequence, recursively calculating in the latent space time-by-time according to differentiable state transition relationships to obtain a first prediction sequence and an uncertainty representation, wherein the first action sequence includes: action data to be evaluated at multiple future times. The first prediction sequence includes: the predicted state of the controlled device at multiple future time points, wherein the uncertainty characterization is used to indicate the degree of prediction dispersion of the predicted state; a first margin is determined based on the uncertainty characterization, and a first deviation is determined based on the degree of deviation of the predicted state in the first prediction sequence from the safety boundary, wherein the position of the safety boundary is adjusted with the first margin; the first deviation is backpropagated to the first action sequence along the recursive path corresponding to the state transition relationship, and the first action sequence is iteratively corrected according to the backpropagation result until the first deviation corresponding to the corrected first action sequence satisfies the convergence condition, thereby obtaining a second action sequence, the second action sequence being used to control the movement of the controlled device.
[0007] Furthermore, the controlled device includes an embodied intelligent robot; the first observation data includes at least one of the joint angles, joint angular velocities, or uniformly coded body state quantities of each joint of the controlled device; the second observation data includes at least one of visual image data, depth data, or environmental perception data; the latent space is a representation space with a dimension lower than that of the second observation data and subject to physical validity constraints.
[0008] Furthermore, the embodied intelligent safety constraint action generation method further includes: extracting a first control frame from the second action sequence, wherein the first control frame is the action data with the earliest timing in the second action sequence; decoupling the first control frame into the target feedforward torque and target reference angular velocity of each joint of the controlled device, and sending them to each first execution component of the controlled device through a first transmission channel, wherein the first transmission channel has a distributed clock synchronization mechanism, the first transmission channel includes a real-time industrial Ethernet bus, and the first execution component includes the servo drive component of each joint; in the next control cycle, based on the reacquired first observation data and second observation data, re-executing the acquisition and determination, the recursion, the determination of the first margin and the first deviation, and the backpropagation and iterative correction.
[0009] Further, the first state representation includes: a first sub-representation and a second sub-representation; the first sub-representation is determined based on the second observation data through a first mapping relationship, the first mapping relationship being used to restrict the features corresponding to the second observation data to a sub-space in the latent space corresponding to the interactive state of the controlled device; the second sub-representation is determined based on the first observation data and a first motion capability parameter, the first motion capability parameter being used to indicate the local sensitivity of the motion mapping of the first execution end of the controlled device in the current configuration, the first execution end being a component in the controlled device used to interact with the environment; the first state representation is a representation obtained by splicing the first sub-representation and the second sub-representation.
[0010] Further, the second sub-representation can be determined as follows: a third sub-representation is determined based on the first observation data, the third sub-representation being used to represent the basic motion state of the controlled device; according to the first weight data, the components of the first motion capability parameter corresponding to different local motion directions are weighted to obtain a fourth sub-representation, the first weight data being used to adjust the influence intensity of different local motion directions in the second sub-representation; the third sub-representation and the fourth sub-representation are added to obtain the second sub-representation; wherein, the first motion capability parameter includes a parameter group composed of the singular values of the pseudo-inverse of the Jacobian matrix of the first execution end in the current configuration of the controlled device, arranged in a predetermined order.
[0011] Furthermore, before the step of recursively calculating the state transition relationship in the latent space according to the first state representation and the first action sequence, the embodied intelligent safety constraint action generation method further includes: projecting the first state representation onto a compliance region according to a first constraint condition to determine a second state representation, wherein the first constraint condition includes multiple safety conditions obtained by local linearization near the first state representation, the compliance region is a region in the latent space composed of states that simultaneously satisfy all the above-mentioned inequalities, and the safety conditions include at least one of: control obstacle conditions, geometric collision boundary conditions, or joint limiting conditions; the step of recursively calculating the state transition relationship in the latent space according to the first state representation and the first action sequence includes: recursively calculating the state transition relationship in the latent space according to the second state representation and the first action sequence, time by time.
[0012] Furthermore, determining the second state representation includes: if the first state representation is located within the compliance area, using the first state representation as the second state representation; if the first state representation is located outside the compliance area, using the state within the compliance area that has the smallest distance from the first state representation as the second state representation.
[0013] Further, taking the state with the smallest distance from the first state representation in the compliant region as the second state representation may include: determining, from the inequalities included in the first constraint conditions, the inequalities not satisfied by the first state representation to obtain a first constraint subset; determining the degree of deviation of the first state representation relative to the boundary corresponding to the first constraint subset; correcting the first state representation along the normal direction of the boundary corresponding to the first constraint subset by an amount corresponding to the degree of deviation, so that the corrected state falls on the boundary corresponding to the first constraint subset, and taking the corrected state as the second state representation.
[0014] Furthermore, the uncertainty representation is also used to indicate the correlation between the components of the predicted state; the uncertainty representation is obtained recursively according to a first propagation relationship, which is used to: propagate the uncertainty representation of the previous moment to the current moment based on the degree of local change of the predicted state and the action data of the previous moment relative to the state in the latent space, according to the state transition relationship, and superimpose a first additional quantity to obtain the uncertainty representation of the current moment; the first additional quantity is related to the action data of the previous moment and is positive semi-definite, and the first additional quantity is used to represent the new prediction error caused by at least one of the following, which cannot be obtained by propagation from the uncertainty representation of the previous moment: action execution error, unmodeled dynamics, contact state change, or residual of the state transition relationship.
[0015] Further, determining the first margin based on the uncertainty characterization includes: determining a first comprehensive index based on the total amount of prediction dispersion indicated by the uncertainty characterization and a first sensitivity, wherein the first sensitivity is used to indicate the degree of response of the predicted state to its corresponding action data, and the first comprehensive index increases with the increase of the total amount of prediction dispersion and increases non-linearly with the increase of the first sensitivity; determining the first margin based on a first rate of change of the first comprehensive index, wherein the first rate of change is used to indicate the degree of change of the first comprehensive index over time.
[0016] Further, determining the first margin based on the first rate of change of the first comprehensive index includes: determining the first margin based on a benchmark margin and the first rate of change, wherein the first margin is equal to the benchmark margin when the first rate of change is zero, the first margin is greater than the benchmark margin when the first rate of change is positive, and the first margin is less than the benchmark margin when the first rate of change is negative; and the first margin monotonically increases with the increase of the first rate of change and converges to an upper bound, wherein the upper bound is a preset multiple of the benchmark margin.
[0017] Furthermore, the first rate of change can be determined as follows: the ratio of the difference between the first comprehensive index corresponding to two adjacent prediction times to the time interval between the two adjacent prediction times is taken as the first rate of change.
[0018] Furthermore, the embodied intelligent safety constraint action generation method further includes: adjusting the boundary positions corresponding to the various inequalities included in the first constraint condition according to the first margin, so that when the first margin increases, the compliance area shrinks relative to the safety boundary, and when the first margin decreases, the compliance area recovers to the range before adjustment.
[0019] Further, the determination of the first deviation includes: extracting a first component group from the predicted state according to a first selection relationship, the first component group being used to characterize the predicted angular velocity and predicted torque of each joint of the controlled device; transforming the first component group according to a first mapping matrix to obtain a first transformation result, the first transformation result being used to characterize the predicted spatial motion and predicted inertia of the first execution end of the controlled device in the global coordinate system, the first mapping matrix being a matrix pre-placed in the feature space corresponding to the latent space to characterize the kinematic relationship of the controlled device body, the first execution end being a component in the controlled device used to interact with the environment; substituting the first transformation result into the inequality relationship corresponding to the safety boundary to obtain a first margin value, the inequality relationship including the quadratic term, linear term and constant term of the first transformation result; determining the first deviation based on the out-of-bounds difference between the first margin value and the first margin; wherein, the determination process of the first deviation is completed in the latent space through algebraic operations and does not depend on the call to an external discrete physics engine.
[0020] Further, determining the first deviation based on the boundary difference between the first margin value and the first margin includes: when the first margin value is less than the first margin, taking the difference between the first margin and the first margin value as the first deviation corresponding to the predicted state; when the first margin value is not less than the first margin, determining the first deviation corresponding to the predicted state as zero; taking the sum of the first deviations corresponding to each predicted state in the first prediction sequence as the first deviation corresponding to the first action sequence, and processing the difference between the first margin and the first margin value according to the first mapping function to obtain the first deviation, wherein the first mapping function is a monotonically increasing function that takes a non-negative value and is differentiable in the entire domain.
[0021] Further, the backpropagation of the first deviation to the first action sequence along the recursive path corresponding to the state transition relationship includes: for a target time in the first prediction sequence where the first deviation is not zero, determining a first correction amount of the first deviation relative to the action data at the time to be corrected in the first action sequence based on the degree of change of the predicted state of the first deviation relative to the target time, the progressive accumulation of the degree of local change of the state transition relationship relative to the state in the latent space at each intermediate time before the target time, and the degree of local change of the state transition relationship relative to the action data at the time to be corrected; summarizing the first correction amounts corresponding to each target time according to the dimension of the action data to obtain the backpropagation result.
[0022] Further, obtaining the second action sequence includes: updating the first action sequence in the opposite direction indicated by the first correction amount by a preset step size; re-executing the recursion based on the updated first action sequence and redetermining the first deviation amount; repeating the update and redetermining until the first deviation amount converges to zero, and taking the converged first action sequence as the second action sequence, the second action sequence belonging to the first action set, the first action set being a set of action sequences that ensure that each predicted state obtained by the recursion does not exceed the safety boundary; wherein, the preset step size is adapted to the spectral radius of the state transition relationship relative to the degree of local change of the state in the latent space.
[0023] Furthermore, the embodied intelligent safety constraint action generation method further includes: parsing a first attribute label of an object in the second observation data, the first attribute label indicating the material category of the object; retrieving a first interaction attribute based on the first attribute label, the first interaction attribute characterizing the interaction mode allowed by the object based on its physical material properties, the first interaction attribute including at least one of stiffness parameter, friction parameter, or failure yield parameter; determining the contact state between the first execution end of the controlled device and the object in the first prediction sequence, the first execution end being a component in the controlled device used to interact with the environment; if contact is determined to have occurred, stopping the determination of the first deviation amount based on the geometric distance between the first execution end and the object, and instead determining the first contact parameter based on the first interaction attribute and the first overlap amount, and if the first contact parameter exceeds a first threshold, accumulating the portion of the first contact parameter exceeding the first threshold to the first deviation amount; wherein, the first overlap amount is used to indicate the predicted intrusion degree of the first execution end relative to the geometric boundary of the object, and the first threshold is the allowable contact force corresponding to the failure yield parameter.
[0024] Furthermore, when the first attribute label indicates that the object belongs to the first material category, the first overlap amount is constrained to zero and treated as a hard constraint, and the degree of change of the first contact parameter over time is limited to within a second threshold, wherein the first material category is a brittle material category.
[0025] Furthermore, when the first attribute label indicates that the object belongs to the second material category, the first overlap amount is allowed to be greater than zero, and the first contact parameter is determined based on the first overlap amount and the first elastic parameter, wherein the first elastic parameter changes with the change of the first overlap amount, and the second material category is a flexible deformable material category.
[0026] Furthermore, the embodied intelligent safety constraint action generation method further includes: monitoring a first monitoring signal, the first monitoring signal including at least one of the communication heartbeat signal between the body control unit of the controlled device and the state assessment node, or the refresh delay of the second observation data; in the case of at least one of the loss of the communication heartbeat signal or the refresh delay being higher than a set threshold, overwriting the calculation basis value of the first comprehensive index with a first preset value, so that the first margin reaches the upper bound and the first deviation corresponding to each non-static predicted state in the first prediction sequence is not zero; and causing the second action sequence to converge to a zero-speed state along the direction of the decrease of the first energy index of the controlled device, the first energy index being used to characterize the total kinetic energy of the controlled device, and the convergence process being constrained by the upper limit of the braking capacity of each joint of the controlled device.
[0027] Another aspect of the present invention provides an embodied intelligent safety constraint action generation system based on world model forward reasoning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements any of the embodied intelligent safety constraint action generation methods based on world model forward reasoning as described above.
[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0029] 1. This invention integrates environmental observation and ontological state into a physically constrained low-dimensional latent state, and injects local sensitivity information of the end-effector motion mapping under the current configuration into it. This allows the deduction to retain ontological maneuverability while stripping away perceptual redundancy, and the prediction results to better match actual motion capabilities. At the same time, by orthogonally projecting the latent state onto a compliant region enclosed by safety conditions before the deduction, it provides a starting point for the entire deduction chain that meets the modeled safety requirements. Furthermore, it only makes minimal corrections to out-of-bounds states, and does not modify in-domain states, thus avoiding information loss caused by forcibly pushing safe states to the boundary.
[0030] 2. This invention constructs a comprehensive uncertainty by combining the degree of prediction dispersion and the degree of action sensitivity, and adjusts the safety margin with its rate of change rather than its absolute value. It expands the margin in advance when the system is rapidly entering an unpredictable region before approaching the dangerous boundary, thus achieving early warning of risks. Furthermore, the degree of violation is obtained entirely by algebraic operations within the latent space, without the need to call an external discrete physics engine, eliminating the overhead of cyclic sampling and intersection testing. Since the deduction path is differentiable, it can be backpropagated along the path to locate the specific out-of-bounds step and out-of-bounds amount, pulling the original intention action back to the safe set with minimal action modification cost, rather than discarding and resampling the whole thing, thus balancing safety and the preservation of task intent.
[0031] 3. Based on the material semantic switching contact constraint, this invention upgrades the binary geometric judgment of whether contact is possible to a continuous physical judgment of how much force is required for safe contact. This allows the same evaluation framework to protect brittle objects from damage while allowing necessary compression to be applied to flexible objects. By overwriting the uncertainty base value when communication or time delay fails, the safety domain shrinks to the limit. The action sequence smoothly converges to zero velocity within the braking capacity constraint along the direction of total kinetic energy decrease. The protective blocking is intrinsically triggered by the existing safety mechanism, without the need for additional independent emergency stop logic. Attached Figure Description
[0032] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:
[0033] Figure 1This is a flowchart of the method provided in Embodiment 1 of the present invention.
[0034] Figure 2 This is a comparison chart of the weighted influence of Jacobi singular value features on latent states provided in Embodiment 1 of the present invention.
[0035] Figure 3 This is a schematic diagram of the information branch structure and dimensional contribution provided in Embodiment 1 of the present invention.
[0036] Figure 4 This is a schematic diagram of the geometric relationship between the physical compliance area of the latent space and the safety orthogonal projection provided in Embodiment 1 of the present invention.
[0037] Figure 5 This is a comparison diagram of the correction magnitude-violation depth relationship and the baseline of the unconditional boundary projection provided in Embodiment 1 of the present invention.
[0038] Figure 6 This is a schematic diagram illustrating the recursive evolution of the latent state prediction covariance provided in Embodiment 1 of the present invention.
[0039] Figure 7 The contact force timing effect diagram for impact rate constraint of brittle material provided in Embodiment 1 of the present invention.
[0040] Figure 8 The index surface diagram of comprehensive uncertainty provided in Embodiment 1 of the present invention.
[0041] Figure 9 This is a schematic diagram illustrating the independent influence of motion sensitivity on overall uncertainty, as provided in Embodiment 1 of the present invention.
[0042] Figure 10 This is a dynamic safety envelope margin adjustment characteristic diagram provided in Embodiment 1 of the present invention.
[0043] Figure 11 This is a schematic diagram of the total violation convergence process and action modification cost provided in Embodiment 1 of the present invention.
[0044] Figure 12 This is a schematic diagram of propagation attenuation provided in Embodiment 1 of the present invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0046] It should be noted that, in this application, the environmental observation matrix can be an observation data tensor output by an RGB camera, depth camera, LiDAR, haptic array, or other sensing device capable of acquiring environmental physical information mounted on the robot body; the environmental observation matrix includes at least: visual image information that can be used to characterize the spatial structure of the environment; additionally, optionally, the environmental observation matrix may also include: depth information, point cloud information, object semantic segmentation layer, or feature layer calculated from multiple basic observation channels. In this application, an embodied robot refers to a robot system that possesses a physical body, can mechanically interact with the real physical environment through joint actuators, and completes tasks such as grasping, moving, assembling, or human-robot collaboration by relying on a perception-decision-execution closed loop; latent space refers to the low-dimensional state representation space formed by compressing the high-dimensional original observations and the body state through feature mapping. In this application, the latent space is also called the physical information latent space because robot kinematics and safety constraint information are explicitly injected.
[0047] Example 1
[0048] This embodiment discloses a method for generating embodied intelligent safety constraint actions based on look-ahead reasoning using a world model. This application addresses the challenges embodied robots typically face when performing tasks in highly unstructured environments: extremely high environmental observation dimensions containing a large amount of visually redundant information irrelevant to the task; numerical sensitivity of end-effector motion mapping when the robot approaches kinematically singular configurations; perception and model errors causing state predictions to fall into unsafe regions; and the difficulty of meeting the low-latency requirements of real-time control due to the overhead of traditional methods relying on external discrete physics engines for collision detection and the cyclic sampling and intersection testing overhead. Furthermore, considering that traditional deterministic world models only output a single future state and cannot quantify the prediction uncertainties introduced by factors such as occlusion, unclear contact states, or unknown material properties, thus failing to tighten safety constraints before the robot actually approaches a dangerous boundary, this method improves the reliability of counterfactual action evaluation and safety action generation for embodied robots under low-latency conditions through the synergistic effects of physical information latent manifold mapping, latent state safety orthogonal projection, mean-covariance synchronous recursion, dynamic safety envelope based on action sensitivity perception, and differentiable safety correction based on time backpropagation. Figure 1 The overall method flowchart of this embodiment is shown. As can be seen from the figure, this embodiment includes the following steps:
[0049] Step 1: Obtain the current physical state vector of the embodied robot and the current environmental observation matrix, and project and map them into an initial latent space state vector with physical constraints.
[0050] The physical state vector refers to a low-dimensional vector composed of the joint angles, angular velocities, or uniformly encoded physical state quantities of the robot's joints, reflecting the robot's configuration and speed of motion. The environmental observation matrix refers to a high-dimensional observation tensor composed of visual images, depth information, or other environmental perception data, reflecting the external environment of the robot.
[0051] The initial latent space state vector refers to the unified low-dimensional state representation obtained by fusing the two types of heterogeneous information mentioned above after feature transformation and manifold constraints. The emphasis on physical constraints is because the latent state is not an arbitrary high-dimensional statistical code, but is explicitly restricted in the physical effective subspace corresponding to the robot's perceptible and interactive state during the construction process, and explicitly injected with the local kinematic maneuverability information of the robot's current configuration.
[0052] In this embodiment, to avoid the problems of needing to process a large amount of visually redundant information that is not directly related to the task when directly extrapolating the future state in the original observation space, and the difficulty in ensuring that the prediction results conform to the robot's local motion capabilities, environmental observation and robot body state are set as two mutually cooperating information branches:
[0053] For the environmental observation branch, the image, depth map or other environmental observation tensors are first expanded into a unified observation vector. Then, the environmental features related to the current task are extracted through linear feature transformation. Subsequently, the environmental features are pulled back into a predefined physical effective subspace through the latent manifold projection operator.
[0054] For the robot body state branch, the basic motion state features are formed by using the current joint angle and joint velocity. At the same time, the pseudo-inverse of the Jacobian matrix of the end effector under the current configuration is calculated, its singular value features are extracted, and different singular directions are weighted by the physical constraint scaling vector, and then added to the basic motion state features.
[0055] Finally, the environmental latent features and the proprioceptive kinematic latent features are combined into a complete latent state through a splicing operation.
[0056] For example, in this embodiment, the latent manifold mapping for physical information injection can be calculated using the following formula:
[0057]
[0058] in, for The initial latent space state vector is obtained by constantly fusing environmental information, robot body state, and local kinematic features. This is a subsurface manifold projection operator used to constrain environmental features to a physically effective subsurface manifold. superior; This is the environmental observation feature transformation matrix, used to map high-dimensional environmental observations to a lower-dimensional environmental feature space; for Environmental observation matrix at any given time; This is a vectorization operation used to convert environmental observation tensors into column vectors; For robots in The ontology state vector at time t. This is the ontology state feature transformation matrix, used to transform... The mapping is to a vector with the same dimension as the Jacobian singular value features; It is a physical constraint scaling vector, each component of which corresponds to a Jacobian singular direction, used to adjust the influence intensity of different local motion directions in the latent state; The pseudo-inverse of the Jacobian matrix for the robot's end effector in the current configuration; This means extracting singular values from a matrix and assembling them into a vector in a predetermined order; For Hadamard product; This is a vector concatenation operation. Figure 2 This diagram illustrates a comparison of the weighted influence of Jacobi singular value features on latent states in this embodiment. Figure 2 In the diagram, (a) is the original Jacobian pseudo-inverse singular value distribution map, and (b) is a schematic diagram of the latent state injection after being weighted by the physical constraint scaling vector. Figure 3 The diagram illustrates the information branching structure and dimensional contribution of the latent manifold mapping in this embodiment.
[0059] Understandably, the purpose of this formula is to simultaneously encode what exists in the environment and how the robot can move in its current configuration into the latent state: the environment observation branch provides external spatial information, the body state branch provides the robot's own motion basis, and the Jacobi singular value branch further provides the sensitivity of local motion mapping in the current configuration; the singular value of the Jacobi pseudo-inverse can reflect the local amplification characteristics and numerical sensitivity of the robot when performing end-effector motion mapping near the current configuration. Injecting this singular value feature into the latent state can allow the latent space to retain the local maneuverability information of the robot's current configuration.
[0060] It should also be noted that directly adding the Jacobian pseudoinverse matrix to the ontology state vector does not hold true in terms of dimension. To enable unified computation between the two, this embodiment extracts the singular value vector of the Jacobian pseudoinverse and scales the vector using physical constraints. Weighting different singular directions retains the technical idea of scaling Jacobian singular values while avoiding mathematical conflicts caused by directly adding matrices and vectors; similarly, concatenation is used... Instead of direct summation, this prevents different information branches from canceling each other out before they are aligned, and allows subsequent models to utilize environmental information and ontological kinematics information respectively.
[0061] It should be noted that the Jacobian pseudoinverse may exhibit large singular values when the robot approaches a kinematically singular configuration. The role of scaling vector is not only to adjust the importance of different motion directions, but also to suppress the excessive amplification of specific singular directions in the latent state; however, relying solely on this scaling vector cannot automatically guarantee numerical stability, and in actual implementation, it is still necessary to ensure that the calculation process of the Jacobi pseudo-inverse has reasonable numerical conditions.
[0062] Step 2: By injecting a control barrier function with physical boundaries, the initial latent space state vector is orthogonally projected onto the physically compliant region defined by safety constraints to obtain the safe latent state.
[0063] The physical compliance region refers to a convex polyhedral safety domain enclosed by several locally linearized safety conditions in the latent space. Latent states falling within this domain satisfy all modeled safety requirements of the robot at the current modeling accuracy. Control barrier functions are scalar functions in control theory used to characterize the invariance of the safety set. They apply inequality conditions to the states to ensure that the system states never exceed a predefined safety set.
[0064] In this embodiment, to prevent the initial latent states caused by perception errors, feature mapping errors, or insufficient training data coverage from falling into areas that do not meet safety requirements, such as a latent state that corresponds to the end effector entering the interior of an obstacle, the joint state exceeding the allowable range, or the distance between the robot and the workspace boundary being less than the predetermined safe distance, thereby contaminating all subsequent look-ahead inferences, it is necessary to perform a safe projection on the latent states before performing autoregressive prediction.
[0065] Specifically, several safety conditions that the robot needs to satisfy (e.g., obstacle control conditions, geometric collision boundaries, joint limit conditions, or other physical constraints) can be converted into affine inequality constraints in the latent space near the current state, with each row of constraints representing a local safety boundary. For latent states that are already within the safety domain, no additional correction is made; for latent states that are outside the safety domain, the safest point closest to the original latent state is found by minimizing the Euclidean distance between the latent states before and after the correction.
[0066] For example, in this embodiment, the safe orthogonal projection of the latent state can be obtained by solving the following quadratic programming problem:
[0067]
[0068] in, This is the initial potential state before any safety corrections have been made; The potential state is yet to be optimized; This is the final safe potential state; Let L be the L2 norm of the vector; This is the latent space safety constraint coefficient matrix, where each row corresponds to a safety condition that has undergone local linearization. This is the safety intercept vector, used to define the boundary positions of each safety constraint.
[0069] When the effective constraint set at the optimal projection point Determined and valid constraint matrix When the row rank is full, the local closed-form solution of this quadratic programming problem can be expressed as:
[0070]
[0071] in, From The matrix composed of the valid constraint rows selected from the list; This is the corresponding safety intercept vector; Used to solve for the Lagrange multipliers required to satisfy effective constraints; This represents the constraint deviation of the original potential state relative to the effective safety boundary. Figure 4 This diagram illustrates the geometric relationship between the latent space physical compliance area and the safety orthogonal projection in this embodiment. Figure 5 The diagram shows a comparison between the correction magnitude-violation depth relationship and the baseline of the unconditional boundary projection in this embodiment.
[0072] Understandably, the physical meaning of this closed-form expression is: first, calculate the degree to which the original latent state violates the effective safety boundary; then, perform a minimum-amplitude correction along the direction spanned by the effective constraint normal vectors, so that the corrected latent state falls exactly on the corresponding safety boundary; if the original latent state already satisfies all inequality constraints, then the optimal solution is... In other words, secure projection does not unconditionally push all latent states to the constraint boundary, but only makes necessary corrections when a latent state violates the security constraint, thereby avoiding information loss caused by forcibly moving the state inside the security domain to the boundary.
[0073] It should be noted that, and This can be understood as a latent space affine safety constraint based on the concept of control barriers, but the above static projection formula itself is not equivalent to a standard CBF controller that includes complete system dynamics; moreover, what this safety projection can guarantee is... Satisfying the latent space safety constraints defined by the current model does not mean that collisions will never occur under any observation error, any model error, or any actuator bias. Rather, it means that subsequent look-ahead inferences will start from a latent state that satisfies the currently modeled safety conditions, thereby reducing the possibility of generating states that obviously violate the safety boundary.
[0074] Step 3: Based on the current candidate action sequence and the safe latent state, perform autoregressive look-ahead state deduction in the latent space, recursively calculate the mean of the predicted state for the next N time steps, and simultaneously propagate the latent state prediction covariance to generate a predicted state sequence.
[0075] Here, the candidate action sequence refers to the sequence of action vectors for the next N time steps to be evaluated, provided by the planner or policy network; autoregressive prospective state deduction refers to a rolling deduction process that starts from the current safe potential state, sequentially inputs candidate actions, and gradually predicts the potential state at the next time step through a state transition function, using the prediction results as the input for the next time step after that. Prediction covariance refers to a symmetric positive semi-definite matrix used to characterize the prediction uncertainty of each potential state component and the correlation between different components.
[0076] In this embodiment, to avoid the shortcomings of outputting deterministic future latent states alone, which cannot reflect the impact of unknown environment, model approximation and observation error on the prediction results, such as the same action may produce a relatively deterministic result in a fully observable environment, but may correspond to multiple different future results in an environment with occlusion, unclear contact state or unknown material properties, this step not only recursively calculates the mean of the future latent states, but also simultaneously recursively calculates the covariance matrix of the latent states.
[0077] For example, in this embodiment, the mean of future latent states can be obtained recursively by the following formula:
[0078]
[0079] At each prediction time, using the current mean and candidate actions as the linearization centers, calculate the Jacobian matrix of the state transition function relative to the latent state:
[0080]
[0081] Furthermore, the latent state prediction covariance can be obtained recursively using the following formula:
[0082]
[0083] in, for The mean of the predicted latent state at time t; From At all times Candidate actions at any given moment; For parameters The latent state transition function, The symbol is for partial differentials; The first-order partial derivative matrix of the state transition function with respect to the latent state at the current predicted mean and candidate action reflects how small deviations in the latent state at the previous time step are propagated to the next time step. for The latent state prediction covariance matrix at time 1; The covariance matrix represents the additional uncertainty associated with the candidate action, used to characterize the new prediction error that cannot be obtained by linear propagation of the covariance from the previous time step, resulting from action execution error, unmodeled dynamics, contact state changes, or state transition model residuals. Figure 6 This diagram illustrates the recursive evolution of the predicted covariance of different candidate action sequences in this embodiment.
[0084] It is understandable that in the recurrence relation The description addresses how existing uncertainties propagate after local nonlinear state transitions: if a certain direction corresponds to a large local gain, then the small uncertainty in that direction at the previous moment will be amplified in subsequent deductions; and The term describes the new uncertainties introduced by the current action and model error. When a candidate action causes the robot to enter an area with complex contact states, significant visual occlusion, or high kinematic sensitivity, the corresponding... The uncertainty is relatively small when the candidate action is located in a region where there are sufficient observations and the model is relatively reliable.
[0085] It should be noted that, It should be a positive semi-definite matrix to ensure that the result obtained through recursion is... It still retains the mathematical properties of the covariance matrix. However, the above covariance recursion is based on the first-order linearization of the state transition function near the mean, which is a local approximation; when When there is strong nonlinearity in the current prediction region, the covariance obtained by first-order propagation may not fully represent the true future state distribution. Therefore, this model is more suitable for short-term rolling simulations, or for re-initializing and correcting future state predictions after each new observation. Furthermore, if If the covariance includes an estimate of insufficient model knowledge, then the covariance may contain a component of cognitive uncertainty.
[0086] Furthermore, in this embodiment, the forward-looking state deduction may also include non-geometric contact safety assessment logic based on interaction availability.
[0087] Interaction availability refers to the set of interaction methods that a robot is allowed to perform based on the physical material properties of an environmental object; for example, a glass can be gently held but not squeezed forcefully, while a sponge can be compressed to a certain extent. Physical material semantic labels refer to the material category identifiers that are labeled for each object in the environmental observation matrix through semantic segmentation or a prior knowledge base.
[0088] In this embodiment, to avoid the traditional collision avoidance penalty based on geometric Euclidean distance from incorrectly classifying all contact as a collision risk in scenarios where contact itself is the task objective (e.g., grasping, pressing, assembling), this step introduces the following deterministic evaluation branch:
[0089] Parse the physical material semantic labels of each object in the environmental observation matrix; based on these labels, retrieve material interaction attributes from a pre-set hardware storage register, including: stiffness coefficient. coefficient of friction and failure yield stress In the predicted state sequence, the physical contact state between the end effector and the object is evaluated. When contact is determined to occur, the collision avoidance penalty calculation based on geometric Euclidean distance is paused, and the contact constraint evaluation based on physical semantics is activated instead.
[0090] For example, the predicted contact internal forces can be approximated using linear elasticity as follows:
[0091] ,
[0092] in, The depth of the end effector intrusion obtained from forward-looking simulation; when Exceed When the allowable contact force is reached, the excess stress difference is converted into an impedance control upper limit penalty term and added to the intrinsic safety boundary violation degree described later.
[0093] Based on this, a deterministic decision branch for switching specific material constraints can also be executed:
[0094] When the material semantic tag indicates that the target object is a highly brittle material (such as glass), the failure yield stress threshold is set to the rigid limit, and geometric state overlap within the latent space is strictly prohibited in violation calculations (i.e., forced). As a hard constraint, the time derivative of the normal contact force is limited to a preset safe impact rate. Within, that is, required When the material semantic tag indicates that the target object is a flexible and deformable material (such as a sponge), the geometric overlap tolerance is enabled in the violation calculation, allowing the predicted coordinates of the end effector to intrude into the original geometric boundary of the object, and the intrusion depth is multiplied by the nonlinear elastic modulus of the sponge. Establish dynamic reaction force constraints, i.e. . Figure 7 The diagram shows the contact force timing effect of the impact rate constraint on brittle materials in this embodiment.
[0095] Understandably, the essence of this branch of logic is to transform the binary geometric problem of whether or not contact is possible into the continuous physical problem of how much force is required for safe contact, so that the same safety assessment framework can both protect brittle objects from damage and allow necessary compression operations to be performed on flexible objects.
[0096] Step 4: Based on the predicted covariance and the sensitivity of the future state to the action, calculate the comprehensive prediction uncertainty, extract the cognitive uncertainty gradient, and dynamically adjust the safety envelope margin accordingly.
[0097] Among them, safety envelope margin refers to the additional protective distance reserved on top of the safety boundary; the larger the value, the greater the reserved distance between the robot and obstacles, joint limits, or other dangerous boundaries. Cognitive uncertainty gradient refers to the rate of change of overall predictive uncertainty over time, reflecting whether the system is entering a more unpredictable region.
[0098] In this embodiment, to overcome the shortcomings of evaluating risk solely based on the magnitude of covariance: two future states may have the same covariance trace, but their sensitivities to motion disturbances differ. For future states with low motion sensitivity, even with some prediction dispersion, small control errors are unlikely to cause drastic state shifts; while for future states with high motion sensitivity, small motion errors may be significantly amplified by the state transition model, causing the robot to rapidly approach or cross the safety boundary; this step combines the overall variance of the latent state with the sensitivity of the future state relative to the motion to construct a comprehensive uncertainty index.
[0099] For example, in this embodiment, the overall uncertainty of the future potential state can be calculated by the following formula:
[0100]
[0101] in, The trace of the covariance matrix represents the sum of the prediction variances of each latent state component; The Jacobian matrix is the mean of the future latent state relative to the previous candidate action, used to describe the degree of response of the mean of the future latent state when the candidate action changes slightly; The Frobenius norm is used to compress the action sensitivity matrix into a single nonnegative scalar. This is the amplification factor for motion sensitivity. Figure 8 The index surface plot of the comprehensive uncertainty in this embodiment is shown; Figure 9This diagram illustrates the independent influence of motion sensitivity on overall uncertainty under the condition of equal covariance traces in this embodiment.
[0102] Understandably, the introduction of the exponential function achieves a nonlinear amplification of the sensitivity to action: when the action sensitivity is low, the overall uncertainty is mainly determined by the covariance; when the action sensitivity increases significantly, even if the covariance has not increased significantly, it can improve the overall uncertainty evaluation result in advance and play the role of "risk early warning".
[0103] After obtaining the overall uncertainty, the next step is to determine whether the uncertainty is rising rapidly. A relatively large one... This indicates that the future state itself is relatively dispersed or highly sensitive to action; a larger... This indicates that the system is rapidly entering a more unpredictable region. For the latter, the safety envelope should be expanded in advance before the robot actually approaches the danger boundary. For example, the dynamic safety envelope margin can be calculated using the following formula:
[0104]
[0105] in, As a baseline safety envelope margin; It is the error function; This is the scaling factor for the rate of change of uncertainty. In the discrete-time implementation, The combined uncertainty quotient between two adjacent prediction times can be used. Approximate, where This represents the time interval between adjacent prediction times. Figure 10 The diagram shows the dynamic safety envelope margin adjustment characteristics of the error function in this embodiment.
[0106] It is understandable that the error function is continuous and monotonic in the real domain, and the output has a finite range. Therefore, while preserving the direction of change, the envelope adjustment amplitude can be limited to avoid the safety envelope margin from increasing infinitely with the growth rate of uncertainty: when hour The reference envelope is used; when hour Expand the safety envelope to maintain a greater reserve distance; when hour Appropriately reduce the extra reserve margin to avoid the system remaining in an overly conservative state for a long time.
[0107] Furthermore, it is also possible to utilize Adjusting the safety intercept vector in step 2 The corresponding boundary position, for example, by taking This allows the safe zone to shrink further relative to the danger boundary when uncertainty increases rapidly, while gradually recovering to the baseline safe zone when uncertainty decreases.
[0108] It should be noted that the inputs to both the exponential function and the error function must be dimensionless quantities: when the action and latent state have been dimensionlessly standardized, , It can be set to a dimensionless coefficient; otherwise It should have dimensions opposite to the motion sensitivity norm. It should have the reciprocal dimension of the rate of change of uncertainty; at the same time, before using the covariance trace, it should be ensured that the components of the latent state have a uniform or comparable scale, otherwise the latent variable with a large numerical scale may be... China occupies an unreasonable dominant position.
[0109] Step 5: Traverse the predicted state sequence to calculate the intrinsic safety boundary violation degree. When the violation degree is detected to exceed the dynamic safety envelope margin, extract the out-of-bounds difference as the violation penalty scalar, perform backpropagation in time (BPTT) along the differentiable state transition path in the latent space to obtain the safety correction gradient, and project the candidate action sequence onto the safety action manifold through iterative gradient descent to generate the target compliant action sequence.
[0110] Among them, the intrinsic safety boundary violation degree refers to the non-negative scalar value of the quantified predicted state exceeding the safety boundary by algebraically calculated entirely within the latent space; it is called intrinsic because its calculation does not depend on any external discrete physics engine. The safety action manifold refers to the set of all action sequences in the action space that enable the look-ahead inference results to satisfy the safety constraints throughout.
[0111] In this embodiment, to eliminate the overhead of cyclic sampling and intersection testing caused by calling an external discrete physics engine, the violation calculation is designed as a pure algebraic mechanism with O(1) time complexity; specifically: directly extract the mean of the predicted latent state at the current time step. The corresponding dimension components of the predicted angular velocity and predicted torque of the robot joints are denoted as the selection matrix. ,Right now The component of this dimension is compared with the Jacobian matrix of the ontology kinematics, which is hard-coded in the feature space. Perform a matrix multiplication to obtain the predicted spatial velocity and predicted inertia tensor of the end effector in the global coordinate system; then substitute them into the quadratic inequality defined by the control obstacle function:
[0112] ,
[0113] in, for The safety margin function at time step 1 is used to characterize the margin of the predicted motion state relative to the dynamic safety boundary. This indicates that the predicted speed is within the permissible safe range and has a certain safety margin; when This indicates that the predicted velocity is exactly within the safety margin; This indicates that the predicted speed exceeds the permissible safety zone, posing a risk of collision, joint overrun, excessive impact, or other violation of safety constraints. Indicating that robots will be in the future The predicted motion velocity at each time step is generally obtained by mapping the predicted joint angular velocity through the robot's Jacobian matrix: in, It predicts the joint angular velocity vector. The kinematic Jacobian matrix is determined based on the robot's current configuration; for example, for a six-DOF end effector. It is usually a six-dimensional spatial velocity, also known as a kinetic spinor or twist: ,in, The linear velocity of the end effector in the global coordinate system; This represents the angular velocity of the end effector in the global coordinate system. This is the coefficient matrix of the quadratic terms; This is the vector of coefficients for the first-order terms; The constant bias term is a scalar term that is independent of the current velocity to be tested. Its main function is to adjust the overall position and scale of the safety boundary.
[0114] Based on this, and combining the dynamic envelope from step 4, the first The penalty scalar for a step violation can be defined as:
[0115]
[0116] That is, when the solution to the CBF inequality falls within the dynamic safety envelope (including cases where the result is negative), its out-of-bounds difference is output as the violation penalty scalar for that step; if material contact evaluation is activated in step 3, the impedance penalty term exceeding the failure yield stress is also accumulated. middle.
[0117] At a certain time step After the scalar of the penalty for violation is non-zero, the state transition equation in the latent space is used. Given its deterministic and differentiable properties, the partial derivative matrix of the scalar with respect to each dimension of the initial candidate action sequence is calculated by backpropagation along time.
[0118] For example, the penalty scalar for violations is relative to the first The partial derivatives of the step action can be expanded using the chain rule as follows:
[0119]
[0120] ,
[0121] in, For from the first Step to the first The state transition Jacobian matrix product of the step has the physical meaning of the first step. After a step-by-step motion perturbation undergoes multiple state evolutions, the first step affects the second step. The cumulative effect of step states; Let be the first-order partial derivative matrix of the state transition function with respect to the action. The safety correction gradient is obtained by summing the partial derivatives of all violation time steps along the action dimension. ,in Total number of violations.
[0122] Subsequently, iterative gradient descent updates are performed on the candidate action sequence in the opposite direction of the safety correction gradient:
[0123]
[0124] in, To update the step size; For each iteration round, after each update, the forward deduction in step 3 is re-executed and the total violation is recalculated until the violation generated by the updated candidate action sequence in the forward deduction converges to zero. The convergence result is established as the target compliant action sequence located on the safe action manifold.
[0125] Understandably, the essence of this mechanism is a counterfactual gradient correction: the world model first answers in the latent space what will happen in the future if the sequence of actions is executed. Once a safety violation is foreseen, it uses a differentiable computation graph to precisely answer how to modify which step of the action and how much to modify it to eliminate the violation. In this way, the original intention action is pulled back to the safe manifold with minimal action modification cost, rather than discarding the entire sequence of actions.
[0126] It should be noted that, due to Since it is not differentiable at zero, a softening approximation (such as the softplus function) can be used in engineering implementation to ensure the numerical stability of BPTT; at the same time, the step size... It should be adapted to the spectral radius of the Jacobian state transition to avoid divergence in the gradient descent process. Figure 11 The diagram illustrates the total violation convergence process and action modification cost under different update step sizes in this embodiment. Figure 12 This diagram illustrates the propagation and attenuation of the BPTT safety correction gradient along the time step in this embodiment.
[0127] Step 6: Fault-tolerant control flow.
[0128] In this embodiment, for extreme scenarios where the embodied robot encounters hardware communication disruptions in a highly unstructured environment, the following fault-tolerant control flow can also be set up:
[0129] Before executing step 1, continuously monitor the communication heartbeat packets between the robot's microcontroller and the cloud-based global state assessment node; when a heartbeat packet is detected to be lost, or the data bus refresh delay of the environmental observation matrix is higher than the set safety threshold (such as network outage or memory overflow causing delay), the controller immediately forces the calculation base value of the cognitive uncertainty gradient to be overwritten as infinity.
[0130] Under this overwrite condition, the input to the error function in step 4 tends to positive infinity, the output saturates to 1, and the dynamic safety envelope margin reaches its theoretical upper limit. The safety region is shrunk to the maximum extent, and almost all non-static predicted states are judged as violations; at this time, the safety correction gradient in step 5 will unconditionally drive the candidate action sequence to reduce the total kinetic energy of the system. The gradient direction converges abruptly and is constrained by physical and dynamic limitations (the braking torque of each joint does not exceed the maximum braking torque). This mechanism forces the target compliant action sequence to smoothly transition to a zero-speed state, triggering an intrinsic protective blockade. This mechanism causes the robot's default behavior in the event of a perception failure to degenerate into a safe stop, without relying on any additional independent emergency stop logic.
[0131] Step 7: Extract the first control frame of the target compliant action sequence and send it to the underlying hardware servo system for execution via the real-time industrial bus.
[0132] Specifically, after generating the target compliance action sequence, the time index is extracted from the sequence. The first step is to control the frame; decouple the control frame into the target feedforward torque and target reference angular velocity of each physical joint; through the real-time operating system (RTOS) deployed inside the robot, the above control quantities are packaged into process data objects (PDOs), and then distributed to the underlying servo drive board of each joint in real time within a time window of hundreds of microseconds using the distributed clock synchronization mechanism of the EtherCAT industrial Ethernet bus.
[0133] It should be noted that executing only the first frame of the sequence, rather than the entire sequence, reflects the idea of rolling time-domain control: in the next control cycle, the system will re-execute steps 1 to 5 based on the latest observations, so that the safety assessment and action correction are always based on the latest environmental information, further reducing the risk caused by the accumulation of model errors with the extrapolation step size.
[0134] In summary, the basic implementation scheme of this embodiment is to construct an embodied intelligent safety action generation system with a physical information latent space world model as the core, deeply integrating kinematic priors and differentiable safety constraints. This scheme first fuses high-dimensional environmental observations with ontological states infused with Jacobi singular value features into an initial latent state with physical constraints through a bi-branch latent manifold mapping; then, through quadratic programming orthogonal projection based on the control obstacle concept, it ensures that the look-ahead inference starts from the latent state that satisfies the modeled safety conditions; based on this, it performs mean-covariance synchronized autoregressive look-ahead inference on candidate action sequences, and optionally superimposes a non-geometric contact safety assessment based on material availability; furthermore, it constructs a comprehensive uncertainty by integrating the covariance trace and the exponential amplification term of action sensitivity, and uses an error function to analyze its rate of change. Smooth compression is used to dynamically adjust the safety envelope margin; then, an O(1) level algebraic mechanism is used to calculate the intrinsic safety boundary violation degree. Once an out-of-bounds violation is detected, BPTT is executed along the differentiable state transition path to extract the safety correction gradient. The candidate action sequence is projected onto the safety action manifold through iterative gradient descent. Finally, the first frame control command of the compliant action is sent to the servo actuator in hard real time via RTOS and EtherCAT bus, and an intrinsic protective blocking fault-tolerant mechanism is used in the communication fault scenario. Thus, the counterfactual safety evaluation, minimum cost correction and hard real-time execution of the action are realized simultaneously under low latency conditions.
[0135] Example 2
[0136] This embodiment discloses an embodied intelligent safety constraint action generation system based on world model forward reasoning.
[0137] Specifically, the embodied intelligent safety constraint action generation system based on world model forward reasoning can be integrated into an electronic device, which can be a robot body embedded controller, edge computing node, server, or other devices. The server can be a single server, a server cluster composed of multiple servers, a deep learning computing node equipped with a graphics processing unit (GPU) accelerator card, or a distributed computing cluster or cloud robot decision-making platform with GPU acceleration capabilities. When the electronic device is running, it can realize the embodied intelligent safety constraint action generation method based on world model forward reasoning of Embodiment 1 of this application.
[0138] In this embodiment, the embodied intelligent safety constraint action generation system based on world model forward reasoning can also be integrated into multiple electronic devices. For example, the state parsing and real-time distribution can be performed by the robot body embedded controller, and forward reasoning and gradient correction can be performed by the edge or cloud computing nodes. Multiple devices work together to realize the embodied intelligent safety constraint action generation method based on world model forward reasoning in Embodiment 1 of this application.
[0139] In this embodiment, the server can also be implemented in the form of a terminal.
[0140] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating embodied intelligent safety constraint actions based on world model forward reasoning, characterized in that, The method for generating embodied intelligent safety constraint actions includes: First observation data and second observation data are acquired, and a first state representation is determined based on the first observation data and second observation data. The first observation data is used to represent the motion state of the controlled device itself, and the second observation data is used to represent the environment in which the controlled device is located. The first state is represented as a fusion of the first observation data and the second observation data in the latent space; Based on the first state representation and the first action sequence, the first prediction sequence and uncertainty representation are obtained by recursively calculating time-by-time in the latent space according to the differentiable state transition relationship. The first action sequence includes: action data to be evaluated at multiple future time points. The first prediction sequence includes: the predicted states of the controlled device at multiple future time points. The uncertainty characterization is used to indicate the degree of dispersion in the predicted state; A first margin is determined based on the uncertainty characterization, and a first deviation is determined based on the degree to which the predicted state in the first prediction sequence crosses the safety boundary, wherein the position of the safety boundary is adjusted with the first margin. Along the recursive path corresponding to the state transition relationship, the first deviation is propagated backward to the first action sequence, and the first action sequence is iteratively corrected according to the backpropagation result until the first deviation corresponding to the corrected first action sequence satisfies the convergence condition, thereby obtaining the second action sequence, which is used to control the movement of the controlled device.
2. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 1, characterized in that, The first state representation includes: a first sub-representation and a second sub-representation; The first sub-representation is determined based on the second observation data through the first mapping relationship. The first mapping relationship is used to restrict the features corresponding to the second observation data to a subspace within the latent space that corresponds to the interactive state of the controlled device; The second sub-characteristic is determined based on the first observation data and the first motion capability parameter. The first motion capability parameter is used to indicate the local sensitivity of the motion mapping of the first execution end in the current configuration of the controlled device, where the first execution end is a component in the controlled device used to interact with the environment. The first state representation is the representation obtained by concatenating the first sub-representation and the second sub-representation.
3. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 1, characterized in that, Before the step of recursively calculating the embodied intelligent safety constraint action generation method in the latent space according to the differentiable state transition relationship based on the first state representation and the first action sequence, the method further includes: Based on the first constraint, the first state representation is projected onto the compliant region to determine the second state representation. The first constraint includes multiple security conditions that are locally linearized in the vicinity of the first state representation, resulting in inequality relationships. The compliance region is the area in the latent space that simultaneously satisfies all of the aforementioned inequality relationships. The safety conditions include at least one of the following: control obstacle conditions, geometric collision boundary conditions, or joint restraint conditions; The step of recursively calculating in the latent space time-by-time based on the first state representation and the first action sequence according to differentiable state transition relationships includes: Based on the second state representation and the first action sequence, the state transition relationship is recursively applied in the latent space time by time.
4. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 3, characterized in that, The determination of the second state representation includes: If the first state representation is located within the compliant area, the first state representation shall be used as the second state representation; If the first state representation is located outside the compliance area, the state with the smallest distance from the first state representation in the compliance area shall be used as the second state representation.
5. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 1, characterized in that, The uncertainty characterization is also used to indicate the correlation between the components of the predicted state; The uncertainty representation is obtained by recursively probing each time step according to the first propagation relationship. The first propagation relationship is used to: propagate the uncertainty representation of the previous time step to the current time step according to the degree of local change of the predicted state and action data of the previous time step relative to the state in the latent space, based on the state transition relationship; and superimpose the first additional quantity to obtain the uncertainty representation of the current time step. The first additional quantity is related to the action data of the previous moment and is positive semi-definite. The first additional quantity is used to characterize the new prediction error caused by at least one of the action execution error, unmodeled dynamics, contact state change, or residual of the state transition relationship, which cannot be obtained by propagation from the uncertainty characterization of the previous moment.
6. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 1, characterized in that, Determining the first margin based on the uncertainty characterization includes: A first comprehensive index is determined based on the total amount of prediction dispersion indicated by the uncertainty characterization and a first sensitivity. The first sensitivity is used to indicate the degree of response of the predicted state to its corresponding action data. The first comprehensive index increases with the increase of the total amount of prediction dispersion and increases nonlinearly with the increase of the first sensitivity. The first margin is determined based on the first rate of change of the first comprehensive index, wherein the first rate of change is used to indicate the degree of change of the first comprehensive index over time.
7. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 1, characterized in that, The determination of the first deviation includes: According to the first selection relationship, a first component group is extracted from the predicted state. The first component group is used to characterize the predicted angular velocity and predicted torque of each joint of the controlled device. The first component group is transformed according to the first mapping matrix to obtain the first transformation result. The first transformation result is used to characterize the predicted spatial motion and predicted inertia of the first execution end of the controlled device in the global coordinate system. The first mapping matrix is a matrix pre-placed in the feature space corresponding to the latent space to characterize the kinematic relationship of the controlled device. The first execution end is a component in the controlled device used to interact with the environment. Substituting the first transformation result into the inequality relation corresponding to the safety boundary, a first margin value is obtained. The inequality relation includes the quadratic term, the linear term, and the constant term of the first transformation result. The first deviation is determined based on the boundary difference between the first margin value and the first margin. The determination of the first deviation is accomplished through algebraic operations in the latent space and does not depend on calls to an external discrete physics engine.
8. The method for generating embodied intelligent safety constraint actions based on world model forward reasoning according to claim 1, characterized in that, The step of propagating the first deviation back to the first action sequence along the recursive path corresponding to the state transition relationship includes: For the target time in the first prediction sequence where the first deviation is not zero, the degree of change of the first deviation relative to the predicted state at the target time is considered. The state transition relationship is accumulated stepwise with respect to the local changes of the state in the latent space at each intermediate time before the target time, and the state transition relationship is determined with respect to the local changes of the action data at the time to be corrected, to determine the first correction amount of the first deviation amount relative to the action data at the time to be corrected in the first action sequence. The first correction values corresponding to each target time are summarized according to the dimension of the action data to obtain the backpropagation result; The process of obtaining the second action sequence includes: Update the first action sequence in the opposite direction to the direction indicated by the first correction amount according to the preset step size; The recursion is re-executed based on the updated first action sequence, and the first deviation is re-determined; Repeat the update and the re-determination until the first deviation converges to zero. Take the first action sequence at the time of convergence as the second action sequence. The second action sequence belongs to the first action set. The first action set is a set of action sequences that ensure that each predicted state obtained by the recursion does not exceed the safety boundary. The preset step size is adapted to the spectral radius of the state transition relationship relative to the degree of local change of the state in the latent space.
9. The method for generating embodied intelligent safety constraint actions based on world model prospective reasoning according to claim 1, characterized in that, The method for generating embodied intelligent safety constraint actions also includes: Parse the first attribute label of the object in the second observation data, where the first attribute label is used to indicate the material category of the object; A first interaction attribute is retrieved based on the first attribute tag. The first interaction attribute is used to characterize the interaction mode allowed by the object based on its physical material properties. The first interaction attribute includes at least one of stiffness parameter, friction parameter or failure yield parameter. In the first prediction sequence, the contact state between the first execution end of the controlled device and the object is determined, wherein the first execution end is a component in the controlled device used to interact with the environment; If contact is determined to have occurred, the determination of the first deviation amount based on the geometric distance between the first execution end and the object is stopped, and the determination of the first contact parameter is based on the first interaction attribute and the first overlap amount. If the first contact parameter exceeds the first threshold, the portion of the first contact parameter that exceeds the first threshold is added to the first deviation amount. Wherein, the first overlap amount is used to indicate the predicted degree of intrusion of the first execution end relative to the geometric boundary of the object, and the first threshold is the allowable contact force corresponding to the failure yield parameter.
10. A system for generating embodied intelligent safety constraint actions based on world model forward reasoning, characterized in that, The embodied intelligent safety constraint action generation system includes: processor; The memory stores a computer program that, when executed by a processor, implements the embodied intelligent safety constraint action generation method based on world model forward reasoning as described in any one of claims 1 to 9.