A somatic intelligent multi-modal perception decision method and system

By employing a closed-loop processing flow of multimodal residual identification and decision correction, the problem of robot motion instability caused by unobservable properties in existing technologies is solved. Dynamic inference and adaptive control of friction characteristics, material properties and internal structure are realized, thereby improving the robot's execution stability and robustness in complex environments.

CN122133090APending Publication Date: 2026-06-02QINGDAO UNIV OF TECH +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO UNIV OF TECH
Filing Date
2026-05-08
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle unobservable implicit properties in complex environments, such as friction coefficients, material elasticity, and internal structures, leading to unstable robot motion execution results. Furthermore, the incremental cost of sensors is high, or machine learning models lack physical interpretation capabilities.

Method used

By constructing a closed-loop processing flow of action prediction, execution feedback, residual identification, latent variable back-inference, and decision correction, multimodal residual data is used for structured analysis and interpretability evaluation to dynamically infer implicit physical properties and adaptively correct decision parameters.

Benefits of technology

It improves the success rate and stability of robots in complex environments, enhances the accuracy and robustness of identifying friction characteristics, material properties and internal structures, and avoids decision-making bias caused by single-mode errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133090A_ABST
    Figure CN122133090A_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent decision-making technology, and more particularly to an embodied intelligent multimodal perception and decision-making method and system. The method includes the following steps: acquiring multimodal data of embodied interaction, and performing action prediction based on the multimodal data of embodied interaction to obtain action prediction data; acquiring execution feedback data, and calculating state deviation based on the action prediction data and execution feedback data to obtain multimodal residual data; identifying residuals in the multimodal residual data to obtain latent variable data; performing latent variable back-estimation based on the latent variable data to obtain latent variable estimation data; and correcting action decisions based on the latent variable estimation data to obtain action decision correction data. This application, by constructing a residual-driven latent variable back-estimation-decision correction closed-loop processing, enables the system to identify unobservable physical attributes from multimodal deviations, thereby improving the accuracy and interpretability of action decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, and in particular to an embodied intelligent multimodal perception and decision-making method and system. Background Technology

[0002] With the development of embodied intelligence technology, the perception and decision-making capabilities of robots in complex environments have become a research focus. Existing technologies typically use multimodal data, including vision, touch, and force feedback, to perceive the state of a target object and directly generate control strategies based on the perception results. However, these methods generally rely on the assumption of "sufficient observable information," meaning that the environmental state can be directly acquired through sensors. In actual interactions, target objects often possess numerous unobservable latent properties, such as friction coefficient, material elasticity, internal structure, and center of gravity distribution. Although these latent variables cannot be directly observed through a single modality, they can still affect the robot's action execution results. For example, during grasping, low-friction targets are prone to slippage, soft materials are prone to excessive deformation, and targets with uneven internal structures may lead to abnormal forces or posture instability. When facing the above problems, existing technologies typically employ two types of methods: one is to improve perception capabilities by increasing the types or precision of sensors, but this approach suffers from high costs and poor adaptability; the other is to directly fit or optimize action results based on machine learning models, but this approach is mostly black-box modeling, lacks the ability to interpret physical properties, and has weak generalization ability when the environment changes. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes an embodied intelligent multimodal perception and decision-making method and system, thereby resolving at least one of the aforementioned technical problems.

[0004] This application provides an embodied intelligence multimodal perception and decision-making method, including the following steps: Acquire multimodal data of embodied interaction, and perform action prediction based on the multimodal data of embodied interaction to obtain action prediction data; Obtain execution feedback data, and calculate state deviation based on action prediction data and execution feedback data to obtain multimodal residual data; Residual identification is performed on the multimodal residual data to obtain latent variable data; latent variable back estimation is performed based on the latent variable data to obtain latent variable estimated data; Action decision correction data is obtained by adjusting the action decision based on the latent variable estimation data.

[0005] This invention constructs a closed-loop processing flow of action prediction—execution feedback—residual identification—latent variable back-inference—decision correction, extending the decision-making process of traditional embodied intelligence, which relies on explicit perception information, to residual-driven latent variable modeling. This enables the system to dynamically infer implicit physical properties such as friction characteristics, material properties, and internal structure even in the presence of unobservable environmental factors. Compared to methods that rely solely on multimodal data fusion, this method improves the accuracy of anomaly source localization through structured analysis and interpretability evaluation of multimodal residuals, and effectively avoids decision bias caused by single-modal errors. Simultaneously, by feeding the latent variable estimation results back into the action decision model, adaptive correction of decision parameters is achieved, giving the system higher stability and robustness in complex interactive environments, thereby improving the success rate and environmental adaptability of the embodied agent in grasping, manipulating, and interactive tasks.

[0006] Optionally, the action prediction includes: Cross-modal state feature extraction is performed based on embodied interaction multimodal data to obtain interaction state feature data; Action association modeling is performed on the interaction state feature data to obtain action association data; Based on the action-related data, the expected state is deduced to obtain the expected state data; Multimodal consistency verification is performed based on the expected state data to obtain action prediction data.

[0007] This invention extracts cross-modal state features from embodied interactive multimodal data and constructs a mechanism for inferring action associations and expected states. This transforms the action prediction process from relying on a single modality or simple fusion results into a structured modeling system oriented towards action-state responses. This invention can uniformly map multi-source information such as visual, tactile, and force feedback to the same state representation space, thereby improving the consistency and completeness of feature representation. Simultaneously, by clarifying the intrinsic relationship between different control commands and environmental responses through action association modeling, it enhances the interpretability of the prediction process. A multimodal consistency verification mechanism is provided, which can effectively identify prediction conflicts between different modalities, dynamically adjust the dominant modality, or correct prediction results, avoiding overall prediction deviations caused by local perceptual errors.

[0008] Optionally, the state deviation calculation includes: The action prediction data and execution feedback data are matched at the feedback time to obtain the time-matching data. State deviation is calculated based on time-matching data to obtain state deviation data, which includes visual state deviation calculation, tactile state deviation calculation, force feedback deviation calculation, and motion trajectory deviation calculation. The state deviation data is processed by residual coupling to obtain multimodal residual data.

[0009] This invention matches the feedback times of action prediction data and execution feedback data, aligning multimodal information from different sources and with different sampling frequencies under a unified time reference, thus avoiding error amplification caused by timing misalignment. By calculating state deviations from multiple dimensions such as vision, touch, force feedback, and motion trajectory, the invention decomposes the originally mixed overall error into structured deviations with clear physical meaning, thereby improving the precision and interpretability of deviation analysis. Through residual coupling and processing of various state deviations, multimodal residual data with a unified cross-modal expression is formed. This not only reflects single-modal anomalies but also reveals the cooperative deviation relationships between multiple modalities, enabling the system to more accurately capture the sources of anomalies in complex interaction processes.

[0010] Optionally, the residual identification includes: Residual mode determination is performed on the multimodal residual data to obtain residual mode data; Based on the residual modal data, cross-modal residual combination is performed to obtain residual combination data; Perform row latent variable type matching on the residual combination data to obtain latent variable type data; The latent variable data is obtained by eliminating the source of the latent variable based on its type.

[0011] This invention determines the residual mode of multimodal residual data and constructs cross-modal residual combination relationships, transforming the originally scattered deviation information into a combination pattern with structural features, thereby achieving a holistic representation of complex interactive anomalies. Matching residual combinations with latent variable types eliminates reliance on a single mode or indicator to determine the source of anomalies; instead, it identifies implicit physical attributes based on multimodal collaborative features, effectively improving the accuracy and stability of latent variable determination. Simultaneously, through residual source elimination processing, spurious residuals caused by sensor noise, timing errors, or occasional disturbances can be removed, avoiding misidentification of latent variables and enhancing the reliability and robustness of the results.

[0012] Optionally, the latent variable back estimation includes: Sensitive residual data are obtained by performing residual sensitivity matching based on latent variable data; Parameter inversion is performed based on the sensitive residual data to obtain the latent variable parameter data; The residual explanatory power is calculated based on the latent variable parameter data to obtain the residual explanatory power data; Based on the residual explanatory power data, parameter convergence screening was performed to obtain the latent variable estimation data.

[0013] This invention introduces a residual sensitivity matching mechanism to screen key residual terms in multimodal residuals that are highly correlated with specific latent variables. This prevents the parameter inversion process from blindly fitting the overall error to focusing on sensitive features with clear physical meaning, thereby improving the targeting and stability of the inversion process. Latent variable parameters are obtained through multi-path parameter inversion, and the explanatory power of different parameter combinations on multimodal residuals is quantitatively evaluated through residual explanatory power calculation, effectively avoiding the problem of overfitting a single parameter to local residuals. Through parameter convergence screening based on explanatory power, consistency and rationality constraints are imposed on candidate parameters, ensuring that the output latent variable estimation results not only cover the main sources of residuals but also maintain cross-modal interpretability consistency.

[0014] Optionally, the parameter inversion includes: The inversion method is selected based on the sensitive residual data, and the inversion method data is obtained; The sensitive residual data are inverted based on the inversion method data to obtain the latent variable parameter data. The inversion process includes at least one of friction parameter inversion, elastic parameter inversion, center of gravity parameter inversion, and structural parameter inversion.

[0015] This invention adaptively selects the corresponding inversion method based on the sensitive residual term, transforming the parameter inversion process from the traditional unified modeling or single optimization path into a multi-path inversion mechanism oriented towards different physical properties, thereby improving the targeting and accuracy of latent variable estimation. Friction parameter inversion is used for slip-type residuals, elastic parameter inversion for deformation-type residuals, center-of-gravity parameter inversion for attitude and moment anomalies, and structural parameter inversion for local abrupt changes. This ensures that all types of residuals can be processed within their corresponding physical interpretation framework, avoiding misjudgment problems caused by mixed modeling of different residuals.

[0016] Optionally, the calculation of the residual explanation degree includes: Based on the latent variable parameter data, the motion operation is reconstructed to obtain motion reconstruction data; Residual calculations are performed based on action reconstruction data and execution feedback data to obtain latent variable residual data; Based on the latent variable residual data and the latent variable data, residual reduction judgment is performed to obtain the residual reduction data; Residual term coverage is performed based on the residual reduction data to obtain residual coverage data; The residual reduced data and residual covered data are used to evaluate the explanatory power of the latent variable parameter data, thus obtaining the residual explanatory power data.

[0017] This invention reconstructs actions based on latent variable parameters, enabling the system to compare and analyze state responses before and after the introduction of latent variables under unified action conditions, thus making the effects of latent variable parameters explicit. By recalculating the latent variable residuals and judging residual reduction, the ability of each latent variable parameter to correct the original multimodal residuals can be directly quantified, avoiding reliance on experience or local fitting for parameter evaluation. Simultaneously, through residual term coverage analysis, the explanatory range of the same latent variable for different modes and types of residuals can be evaluated, improving the comprehensiveness and rationality of parameter selection. Combining residual reduction and coverage for explanatory power evaluation transforms the judgment of the merits of latent variable parameters from a single error index to a multi-dimensional comprehensive evaluation, effectively improving the accuracy and stability of parameter selection.

[0018] Optionally, the latent variable back estimation includes: Spatial mapping is performed on the latent variable data to obtain spatially mapped data; The spatial mapping data is reconstructed using a probability distribution to obtain the reconstructed distribution data. Perform correlation reasoning on the distributed reconstructed data to obtain correlation reasoning data; By performing a reverse iteration of change optimization based on the correlation inference data, we obtain the latent variable estimation data.

[0019] This invention maps discrete latent variable data to a unified spatial representation domain, enabling latent variables from different sources and of different types to be structurally represented in the same parameter space, thus enhancing the comparability and correlation between latent variables. Modeling the uncertainty of latent variables through probability distribution reconstruction effectively represents ambiguity and noise effects, avoiding instability issues caused by single-point estimation. Combining an association inference mechanism to analyze the dependencies between different latent variables allows the system to identify collaborative or constraint relationships, thereby improving the rationality of the overall inference results. Through a reverse iterative process of variation optimization, the latent variable estimation results gradually converge in multiple rounds of feedback, improving estimation accuracy and enhancing the stability and robustness of the results, providing reliable support for decision-making in complex interactive environments.

[0020] Optionally, the action decision correction includes: Decision analysis is performed on the latent variable estimation data to obtain decision analysis data; Based on the decision impact data, the action risk is reassessed to obtain action risk data; Based on the action risk data, the correction target is determined, and the action correction target data is obtained; The motion parameters are adjusted based on the motion correction target data to obtain the motion correction data; The action decision is updated based on the action correction data to obtain action decision correction data.

[0021] This invention utilizes latent variable estimation data for decision analysis, transforming previously unobservable environmental attributes (such as friction characteristics, material stiffness, and center of gravity distribution) into structured constraints that can be used for control decisions. This allows action generation to no longer rely on single sensory results but reflect real physical interaction conditions. Through action risk reassessment, key risks such as slippage, pressure loss, attitude instability, and trajectory deviation are requantified, avoiding the use of unsuitable original strategies in complex environments. By combining target determination and action parameter adjustment, the system can dynamically optimize key parameters such as clamping force, contact method, execution path, and motion speed for different latent variable scenarios. Closed-loop control is formed through action decision updates, enabling adaptive correction of decision results before execution, improving task success rate, stability, and environmental adaptability.

[0022] Optionally, this application also provides an embodied intelligent multimodal perception and decision-making system for executing the embodied intelligent multimodal perception and decision-making method described above, the embodied intelligent multimodal perception and decision-making system comprising: The action prediction module is used to acquire multimodal data of embodied interaction and perform action prediction based on the multimodal data of embodied interaction to obtain action prediction data; The multimodal residual calculation module is used to acquire execution feedback data and calculate state deviation based on action prediction data and execution feedback data to obtain multimodal residual data. The latent variable estimation module is used to identify residuals in multimodal residual data to obtain latent variable data; and to perform latent variable back-estimation based on the latent variable data to obtain latent variable estimated data. The action decision correction module is used to correct action decisions based on latent variable estimation data, and obtain action decision correction data.

[0023] In summary, this invention predicts actions from multiple sources, including visual, tactile, and force feedback data, establishing a forward mapping relationship between actions and state responses. By performing state deviation calculations with execution feedback and temporal alignment, the overall error is decomposed into multimodal residuals with clear physical meaning, achieving a structured representation of anomalies. Through residual identification and cross-modal combination analysis, complex deviations are mapped into latent variables such as potential frictional characteristics, material properties, or structural differences. Reliable estimation of these latent variables is achieved through parameter inversion and interpretability evaluation mechanisms. The latent variable estimation results are introduced into the action decision correction stage, allowing for targeted adjustments to clamping force, contact strategy, and execution path, forming adaptive control based on real environmental attributes. This invention improves the consistency and accuracy of perception and decision-making, enhancing the system's stability and robustness in interactive environments. Attached Figure Description

[0024] Other features, objects, and advantages of this application will become more apparent from the following detailed description of the non-limiting embodiments, taken with reference to the accompanying drawings: Figure 1 A flowchart illustrating the steps of an embodiment of an embodied intelligence multimodal perception and decision-making method is shown. Figure 2 A flowchart illustrating the steps of an action prediction method according to one embodiment is shown. Figure 3 A flowchart illustrating the steps of a state deviation calculation method according to an embodiment is shown. Figure 4 A flowchart illustrating the steps of a latent variable extraction method according to an embodiment is shown; Figure 5 A flowchart illustrating the steps of an action decision correction method according to one embodiment is shown. Detailed Implementation

[0025] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0026] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. Functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0027] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0028] In an automated assembly line, a robotic arm grips a 50mm diameter rubber part with a Shore A30 hardness. The system acquires multimodal data using an RGB-D camera (30Hz), a tactile array (100Hz), and a six-dimensional force sensor (200Hz). Based on motion correlation, it infers motion prediction data with a predicted pose error of less than ±2mm and a predicted pressure distribution error of less than 10%, enabling early assessment of the contact state. During execution, the system collects feedback data and calculates state deviations to obtain multimodal residual data. The maximum pressure center offset is 4mm, and the peak value of the tangential force residual is 8N, accounting for 20% of the normal force (40N). The system detects a slip displacement residual growth rate exceeding 15% within three consecutive sampling periods, triggering anomaly identification. Through residual identification and back-estimation, the system identifies this as a low-friction latent variable and corrects the friction coefficient from the initial estimate of 0.6 to the range of 0.32–0.38. Based on the residual explained value calculation, the overall residuals decreased by approximately 45% after correction, with the slip-related residuals decreasing by over 60%, yielding high-confidence latent variable estimation results. During the action decision correction phase, the system increased the clamping force from 30N to 42N (not exceeding the 50N upper limit), reduced the transfer acceleration from 1.2m / s² to 0.6m / s², and adjusted the gripping angle by 5° to increase the contact area. Verification showed that the corrected slip displacement decreased from 6mm to 1.5mm, and the contact stability index improved by approximately 70%. In 100 consecutive gripping tests, the system increased the task success rate from 82% to 96%, reduced the slip failure rate by approximately 75%, and kept the pressure loss rate below 2%.

[0029] Please see Figures 1 to 5 This application provides an embodied intelligence multimodal perception and decision-making method, comprising the following steps: S1. Obtain multimodal data of embodied interaction, and perform action prediction based on the multimodal data of embodied interaction to obtain action prediction data; In one embodiment, the system acquires visual images, depth maps, tactile arrays, six-dimensional force sensing data, joint angles, end effector pose, and motion control commands during the robotic arm's grasping, pushing, or handling actions, and aligns them with a uniform sampling period. The system extracts the target pose, contour boundaries, and contact area from the visual data. For example, the system segments the target object based on depth maps or RGB-D data to obtain the target region. It then extracts the contour boundaries by performing edge detection on the target region (e.g., gradient changes exceeding a set threshold). Combining camera intrinsic parameters and depth information, the image coordinates are converted to a three-dimensional coordinate system to obtain the target object's position and pose in the world coordinate system (i.e., target pose). The contact area is determined by identifying regions where the distance between the end effector and the target object is less than a preset threshold (e.g., less than 5 mm). The system extracts the pressure center, contact area, and pressure gradient from the tactile data. For example, each unit in the tactile array corresponds to a pressure value. The system selects tactile units with pressure values ​​greater than a set threshold (e.g., greater than 10% of the maximum pressure) as contact units. The pressure center is determined by the position of each tactile unit and its corresponding contact area. The pressure value is calculated by weighted average. The contact area is obtained by counting the number of effective tactile units and multiplying it by the unit area. The pressure gradient is calculated by the pressure difference between adjacent tactile units and is used to characterize the direction of pressure distribution change. Normal force, tangential force and torque change are extracted from the force feedback data. For example, the system decomposes the triaxial force (Fx, Fy, Fz) and triaxial torque (Mx, My, Mz) output by the six-dimensional force sensor. With the normal direction of the contact surface as a reference, the force is decomposed into normal force Fn (the component along the contact normal direction) and tangential force Ft (the component perpendicular to the normal direction). The torque change is calculated by the torque difference ΔM=M(t+1)-M(t) within the continuous time window (ΔM is the torque change, M(t+1) is the actual torque value at the current sampling time, and M(t) is the actual torque value at the previous sampling time). The system establishes an action association relationship between the above features and the current action command, such as "gripper closure amount - pressure change", "end-effector propulsion direction - displacement trend", and "force application direction - rotation trend". Within a continuous time window, it records the correspondence between changes in action parameters (such as gripper closure amount Δg, end-effector displacement direction Δx, and force application direction Δf) and changes in state characteristics (such as pressure change ΔP, displacement change ΔX, and rotation change Δθ). When a certain action parameter changes, if the corresponding state characteristic continues to change in the same direction in subsequent time windows (for example, the change direction is consistent within 3 consecutive sampling periods), then an association relationship is established.The system infers the expected state after an action is executed based on action correlations (how to infer the expected state after an action is executed based on action correlations), generating action prediction data. For example, when the system infers the expected state after an action is executed based on action correlations, it first uses the target pose, contact state, pressure distribution, force state, and end-effector trajectory at the current sampling time as the initial state. Then, it reads action parameters such as gripper closure amount, end-effector movement direction, applied force magnitude, applied force direction, and execution speed from the current action command. The system determines the state items corresponding to the influence of each action parameter based on the established action correlations. For example, the gripper closure amount is used to infer changes in contact pressure and contact area, the end-effector propulsion direction is used to infer changes in target displacement direction and trajectory, and the applied force direction and contact point position are used to infer rotational trends and attitude changes. For each state item, the system updates it by adding the state increment caused by the action parameters to the current state value to obtain the expected state at the next moment. The action prediction data includes at least predicted pose, predicted trajectory, predicted contact state, predicted slip trend, predicted deformation, and predicted stability. The predicted contact state includes predicted pressure, predicted contact area, and predicted pressure center.

[0030] S2. Obtain execution feedback data, and calculate state deviation based on action prediction data and execution feedback data to obtain multimodal residual data; In one embodiment, the system collects actual feedback data after the action is executed, including actual visual pose, actual tactile pressure distribution, actual force feedback curve, and actual end-effector trajectory. The system matches the action prediction data with the execution feedback data according to the same action stage and time window, for example, matching the predicted pressure in the contact stage with the actual pressure in the contact stage, and matching the predicted pose in the transfer stage with the actual pose in the transfer stage. The system calculates the visual residual, tactile residual, force feedback residual, and trajectory residual respectively. For example, the visual residual can be represented as the deviation between the predicted pose and the actual pose; the tactile residual can be represented as the offset between the predicted pressure center and the actual pressure center; the force feedback residual can be represented as the difference between the predicted tangential force (calculated based on the predicted pressure and the predicted contact state) and the actual tangential force (calculated based on the actual tactile pressure distribution and the actual force feedback curve); and the trajectory residual can be represented as the distance deviation between the predicted end-effector trajectory and the actual end-effector trajectory. The system organizes various residuals according to their occurrence stage, direction, magnitude, and duration to obtain multimodal residual data.

[0031] S3. Perform residual identification on the multimodal residual data to obtain latent variable data; perform latent variable back estimation based on the latent variable data to obtain latent variable estimation data; In one embodiment, the system performs residual mode determination on multimodal residual data, identifying that the current anomaly mainly originates from vision, touch, force feedback, or motion trajectory. If a certain mode satisfies the condition that the residual change value is greater than a preset threshold and the sign of the rate of change remains consistent within three consecutive sampling periods, it is considered a preliminary anomalous mode. Among the preliminary anomalous modes, the one with the largest average residual is selected. If the residuals are of the same maximum value, the mode with the longest residual duration is selected as the dominant mode / abnormal mode. The system performs cross-modal residual combination analysis, forming residual combinations based on the correspondence between multiple residuals. For example, when the visual displacement residual is small (e.g., the residual accounts for less than 5% of the execution feedback data), but the tactile pressure center continuously shifts within 2-3 preset time windows and the tangential force residual increases within 2-3 preset time windows, it is determined to be a latent slip residual combination; when the visual deformation residual increases within 2-3 preset time windows and the contact area expands within 2-3 preset time windows, it is determined to be a soft deformation residual combination; when the visual posture change is not obvious (less than the preset visual posture change threshold) but the torque residual changes abruptly (greater than the preset torque residual threshold), it is determined to be a center of gravity shift or internal structure anomaly residual combination. The system matches the residual combination with a preset latent variable type library to obtain latent variable type data. The preset latent variable type library includes a set of latent variable types predefined based on multimodal residual combinations. The latent variable types include at least one or more of frictional characteristics, material elasticity, center of gravity distribution, and internal structure anomalies. The system first performs a persistence check. If the residual appears only at a single sampling point or lasts for less than three sampling periods, it is identified as transient noise and discarded. Next, a time alignment check is performed. If a modal residual has a fixed time offset relative to other modalities, and can be aligned after delay compensation, it is marked as a synchronization error; otherwise, it is discarded. For visual residuals, if they are accompanied by missing contours or depth data, and there are no corresponding changes in other modalities, they are identified as occlusion pseudo-residuals. Based on the above, the system obtains latent variable data, including latent variable type data, corresponding residual combination data, dominant modality, data of the action phase, residual amplitude data, and residual duration data.

[0032] The system determines the type of latent variable to be inferred based on the latent variable data and selects the corresponding inversion method. If the latent variable type data is friction characteristics, the friction parameter range is inferred based on the tactile pressure center shift, tangential force residual increase, and visual displacement residual proportion in the latent slip residual combination. If the latent variable type data is material elasticity, the elastic parameter range is inferred based on the visual deformation residual, contact area expansion amplitude, and residual duration in the soft deformation residual combination. If the latent variable type data is center of gravity distribution or internal structural anomaly, the center of gravity shift direction or structural anomaly degree is inferred based on the torque residual abrupt change, visual posture change amplitude, and the stage of action. The system compares the candidate parameters obtained by inversion with the original residual combination. If the candidate parameter can explain the dominant mode residual and reduce the corresponding residual amplitude, it is retained as latent variable estimation data. The latent variable estimation data includes latent variable type, parameter estimation range, corresponding residual combination, dominant mode, and stage of action. The back-calculation process is based on the correspondence between residual combinations and physical properties to perform parameter range inversion. Specifically, the system converges the latent variable parameters within a range based on the changing trend, amplitude, and duration of the corresponding residual terms. For example, in friction characteristic back-calculation, the upper limit of the friction coefficient is determined based on the relationship between the tangential force residual and the contact pressure, and the parameter range is narrowed by combining this with the slip duration. In material elasticity back-calculation, the range of elastic parameter changes is determined based on the deformation residual and the expansion amplitude of the contact area. In center of gravity distribution or internal structural anomaly back-calculation, the offset direction or degree of anomaly is determined based on the direction of torque residual change and the amplitude of visual posture change. The system generates candidate parameters within a preset parameter range and gradually filters parameter ranges that meet the residual interpretation conditions by matching and comparing them with the original residual combinations.

[0033] S4. Based on the latent variable estimation data, the action decision is corrected to obtain the action decision corrected data.

[0034] In one embodiment, the system performs decision analysis on the current action strategy based on the types of latent variables and the parameter estimation ranges in the latent variable estimation data, identifying the direction of influence of each latent variable on the action execution. For example, when the latent variable type is friction characteristics and the estimated friction coefficient is lower than a preset friction coefficient threshold, it is determined that the current action has a slip risk; when the latent variable type is material elasticity and the elastic parameter is lower than a preset elastic parameter threshold, it is determined that there is an excessive deformation risk; when the latent variable type is center of gravity distribution or internal structural anomaly, it is determined that there is a risk of attitude instability or abnormal local force. Abnormalities in center of gravity distribution or internal structure are determined by a joint assessment of torque residuals, visual posture changes, and tactile force distribution. For example, under symmetrical clamping or expected linear movement conditions, if the torque residual is greater than a preset torque residual threshold (e.g., 1.5 times the historical average torque) for three consecutive sampling periods, and the direction of visual posture change is consistent with the direction of torque change, an abnormality in center of gravity shift is determined. When the visual posture change is small (e.g., the posture change amplitude is less than 5% of the preset posture threshold), but the tactile pressure distribution is significantly uneven (e.g., the pressure center shift exceeds 10% of the contact area size), and there are abrupt or discontinuous changes in force feedback, an abnormality in internal structure is determined.

[0035] The system reassesses the risks of motion parameters based on the types of latent variables and their corresponding residual combinations. Specifically, it matches the current motion parameters (such as gripper closure, end effector direction, force direction, and motion speed) with the range of latent variable parameters to determine whether the current motion is within a safe range. For example, if the contact pressure corresponding to the current clamping force is insufficient to meet the frictional constraints obtained by reverse calculation, it is marked as slippage risk; if the current applied force exceeds the deformation tolerance range corresponding to the elastic parameter, it is marked as pressure loss risk.

[0036] The system determines the action correction target based on the type of risk. For example, for slippage risk, the correction target is to improve contact stability; for deformation risk, the correction target is to reduce local pressure; and for center of gravity shift risk, the correction target is to adjust the force balance. The system adjusts the motion parameters according to the correction target, such as increasing the gripper closure to increase contact pressure, adjusting the end effector direction to reduce tangential force, reducing the movement speed to reduce impact, or adjusting the gripping point position according to the direction of center of gravity shift. The contact stability refers to the ability of the end effector to maintain a stable contact position, force state, and relative motion state with the target object during contact. Contact stability is characterized by the offset of the pressure center and the change in tangential force: when the offset distance of the pressure center relative to the geometric center of the contact area is less than a preset offset threshold (e.g., less than 10% of the feature size of the contact area), and the change in tangential force is less than a preset force change threshold (e.g., less than 1.2 times the historical average) within 2-3 consecutive sampling periods, the contact stability is considered good; when the offset of the pressure center continues to increase and exceeds the preset offset threshold, and the change in tangential force continues to increase and exceeds the preset threshold (e.g., more than 1.5 times the historical average) within 2-3 consecutive sampling periods, the contact stability is considered to have decreased, and is therefore considered poor.

[0037] The system substitutes the adjusted motion parameters into the motion prediction process, regenerates the predicted state, and compares it with the original residual combination. If the adjustment reduces the dominant mode residual without introducing new abnormal residuals, the correction is deemed effective; otherwise, iterative adjustments are continued within the parameter range. Motion parameters that meet the residual reduction condition and conform to the constraints of the current motion stage are output as motion decision correction data. This data includes the corrected motion parameters, corresponding latent variable types, applicable motion stage, and correction flags.

[0038] Optionally, the action prediction includes: S11. Based on the multimodal data of embodied interaction, perform cross-modal state feature extraction to obtain interaction state feature data; In one embodiment, the system first timestamps and aligns visual images, depth maps, haptic arrays, six-dimensional force sensing data, joint angles, end-effector poses, and motion control commands, mapping data with different sampling frequencies to a unified motion cycle. From the visual data, the system extracts the target object's center coordinates, pose angles, contour boundaries, contact area, and displacement of adjacent frames; from the haptic array, it extracts the pressure center, maximum pressure point, contact area, and pressure distribution gradient (calculated from the pressure difference between adjacent haptic units); from the force feedback data, it extracts normal force, tangential force, torque, and their rate of change; and from the joint states, it extracts end-effector position, velocity, acceleration, and pose changes. These features are then encapsulated within the same time window to obtain interactive state feature data.

[0039] S12. Perform action association modeling on the interaction state feature data to obtain action association data; In one embodiment, the system establishes a motion association model based on the response relationship between motion control commands and interactive state characteristics. The system correlates the gripper closure amount with tactile pressure changes, the end effector propulsion direction with the target object's displacement direction, the applied force direction with the object's rotational trend, and the end effector velocity with the contact impact amplitude / instantaneous force feedback change. If, after a change in a certain motion command within a continuous time window, the target pose, pressure center, or force feedback curve undergoes a stable change (i.e., the change direction is consistent within three consecutive sampling periods, and the change amplitude exceeds a preset threshold), then a correlation is established between the motion command and the state response. The motion association data includes at least the motion type, motion parameters, associated state characteristics, motion direction (if it is a non-mechanical command, the direction is determined by the rate of change), and the motion occurrence stage.

[0040] S13. Based on the action-related data, perform expected state deduction to obtain expected state data; In one embodiment, the system uses the current interaction state as initial conditions and extrapolates the theoretical response after the target action is executed based on action-related data. For example, when the gripper closure amount increases, it extrapolates the change in the contact pressure center, the expansion of the contact area, and the gripping stability; when the end effector advances along a specified direction, it extrapolates the displacement direction, rotation angle, and slippage trend of the target object; when the applied force direction deviates from the target center, it extrapolates the change in the object's posture. The system represents the expected state as follows: ,in, This is the current interaction state. For the current action parameters, This is the action response mapping function, and the state update function constructed based on action correlation. The system obtains the expected pose, expected trajectory, expected contact state, expected pressure distribution, expected slip trend, and expected stability.

[0041] S14. Perform multimodal consistency verification based on the expected state data to obtain action prediction data.

[0042] In one embodiment, the system performs consistency verification on the visual prediction results, tactile prediction results, force feedback prediction results, and trajectory prediction results in the expected state data. If the displacement direction predicted by the vision is consistent with the trajectory prediction direction, and the direction of change of the tactile pressure center is consistent with the trend of change of the force feedback, i.e., the direction signs are the same and remain unchanged for 2-3 consecutive sampling periods, then the prediction results are determined to be consistent. If the visually displayed object is stable, i.e., the visual pose change amplitude is less than a preset pose change threshold (e.g., less than 5%), but the tactile pressure center continues to shift or the tangential force changes abnormally, such as the tangential force change amplitude being greater than a preset force change threshold (e.g., greater than 1.5 times the historical average), then a modal conflict marker is generated. The system determines the dominant modality according to the action stage; for example, before contact, visual and depth data are dominant, and after contact, tactile and force feedback data are dominant. The system encapsulates the expected state after consistency verification, modal conflict markers, prediction confidence, and corresponding multimodal data into action prediction data. The prediction reliability is determined based on the number of modalities participating in the consistency verification and the proportion of consistent modalities. For example, the system uses visual, tactile, force feedback, and trajectory prediction results as modalities participating in the verification, and counts the number of modalities that meet the consistency conditions. For instance, if there are four modalities participating in the verification, and the visual prediction direction, trajectory prediction direction, and force feedback change direction all meet the consistency conditions, while the tactile pressure center change direction does not, then the number of consistent modalities is 3, the total number of participating modalities is 4, and the prediction reliability can be determined to be 75%.

[0043] Optionally, the state deviation calculation includes: S21. Perform feedback timing matching on the action prediction data and execution feedback data to obtain timing matching data; In one embodiment, the system maps motion prediction data and execution feedback data to the same motion time axis based on the triggering time of the motion control command. For the expected pose, expected contact state, expected pressure distribution, and expected trajectory in the prediction data, the system divides the motion stage into approach, contact, clamping, moving, or releasing windows. For the visual feedback, tactile feedback, force feedback, and joint feedback in the execution feedback data, the system performs corresponding matching based on the timestamp and motion stage label. If a certain feedback data has a sampling delay, it is moved forward or backward to the corresponding window according to the sensor delay compensation value, which is determined based on the average historical time offset of each sensor. The system obtains the prediction-feedback correspondence at the same time and in the same motion stage, obtaining time-matched data. The same motion stage refers to the prediction data and feedback data being in the same stage label and the time difference being less than a preset time threshold (e.g., less than one sampling period).

[0044] S22. Calculate the state deviation based on the time-matching data to obtain the state deviation data. The state deviation calculation includes visual state deviation calculation, tactile state deviation calculation, force feedback deviation calculation, and motion trajectory deviation calculation. In one embodiment, the system calculates the state deviations under different modalities. Visual state deviation is determined based on the translation distance, rotation angle difference, and contour boundary offset between the predicted and actual poses; tactile state deviation is determined based on the offset between the predicted and actual pressure centers, the difference between the predicted and actual contact areas, and pressure gradient changes; force feedback deviation is determined based on the differences between the predicted normal force, tangential force, torque, and actual force feedback; and motion trajectory deviation is determined based on the positional distance, velocity difference, and attitude change difference between the predicted and actual end-effector trajectories. This can be expressed as: ,in, Let m be the deviation of mode m at time t. This is the actual feedback status. This is for predicting the state.

[0045] S23. Perform residual coupling processing on the state deviation data to obtain multimodal residual data.

[0046] In one embodiment, the system uniformly organizes the state deviations of visual, tactile, force feedback, and motion trajectory according to the action stage, occurrence time, residual direction, residual amplitude, and duration. If multiple modal residuals occur simultaneously within the same time window and point to the same interaction anomaly, i.e., the residual change direction is consistent and the change sign is the same, then a residual coupling relationship is established. For example, if visual displacement deviation and tactile pressure center shift occur simultaneously, accompanied by an increase in tangential force deviation, and if this increase occurs within 2-3 consecutive sampling periods, the system organizes them into a slip-related residual group; if the tactile contact area abnormally expands and is accompanied by visual deformation deviation, then it is organized into a deformation-related residual group. Multimodal residual data is obtained, which includes residual type, source mode, occurrence stage, residual amplitude, duration, and cross-modal coupling relationship. The residual types include slip-related residuals and deformation-related residuals.

[0047] Optionally, the residual identification includes: S31. Perform residual mode determination on the multimodal residual data to obtain residual mode data; In one embodiment, the system performs modal determination on multimodal residual data according to the residual source mode, occurrence stage, residual amplitude, and duration. If visual pose deviation, contour offset, or target area drift continuously exceeds a preset window (e.g., 2-3 consecutive sampling periods), it is marked as a visually dominant residual; if pressure center offset, abnormal contact area (i.e., exceeding a preset threshold), or sudden pressure gradient change (i.e., the change between adjacent time points is greater than the threshold) occurs continuously, it is marked as a tactilely dominant residual; if the changes in normal force, tangential force, or torque deviate significantly from the predicted value, i.e., the deviation amplitude is greater than the threshold, it is marked as a force feedback dominant residual; if the end trajectory, velocity, or posture is inconsistent with the predicted trajectory, it is marked as a motion trajectory dominant residual. The system obtains residual modal data, which includes at least the dominant mode, auxiliary mode, residual direction, residual intensity, and occurrence stage.

[0048] S32. Perform cross-modal residual combination based on the residual modal data to obtain residual combination data; In one embodiment, the system combines residuals that are causally related (i.e., the order of change is temporally sequential and the direction of change is consistent) within the same action phase and adjacent time windows. If visual displacement deviation, tactile pressure center shift, and tangential force residual occur simultaneously within the same time window (e.g., 2-3 sampling periods), and the residual directions all point towards relative sliding of the contact surface (i.e., the residual change direction is consistent and consistent with the tangential direction of the contact surface), then the combination is a slip-type residual. If visual deformation increases, tactile contact area expands, and normal force residual decreases (i.e., the three change monotonically within 2-3 consecutive sampling periods) occur simultaneously, then the combination is a soft deformation type residual. If torque residual abruptly changes, visual rotation deviation, and trajectory attitude shift occur simultaneously, then the combination is a center of gravity shift type residual. The system records the combined residual type, participating mode, occurrence stage, and combination relationship as residual combination data.

[0049] S33. Perform row latent variable type matching on the residual combination data to obtain latent variable type data; In one embodiment, the system matches the residual combination data with a preset latent variable type library, which is a predefined set of correspondences between residual combinations and latent variable types. If the residual combination shows an increasing trend in slip displacement over 2-3 consecutive sampling periods, tangential force less than a preset threshold, and continuously decreasing contact stability, it is matched as a friction coefficient latent variable; if the residual combination shows an increase in contact area exceeding a preset proportion (e.g., exceeding 10%) over 2-3 consecutive sampling periods, increased deformation, and decreased pressure distribution gradient, it is matched as an elastic modulus latent variable; if the residual combination shows a torque change greater than a preset threshold, rotation angle deviation exceeding a preset angle threshold (e.g., greater than 5°), and attitude change direction deviating from the expected direction exceeding a preset angle threshold, it is matched as a center of gravity offset latent variable; if the residual combination shows a sudden change in local pressure, discontinuous changes in deformation between adjacent regions, and a force response distribution deviation exceeding a threshold, it is matched as an internal structural anomaly latent variable. The system obtains latent variable type data.

[0050] S34. Perform residual source exclusion processing based on the latent variable type to obtain latent variable data.

[0051] In one embodiment, the system performs exclusion verification on the sources of residuals corresponding to latent variable type data. If the residual appears only at a single sampling point and does not appear continuously in adjacent time windows, it is determined to be transient noise and discarded; if the residual exists only in the visual modality and is accompanied by occlusion, reflection, or target detection box jumps, it is determined to be a visual spurious residual; if the tactile or force feedback residual completely coincides with the control command switching time and does not reappear in the subsequent 2-3 sampling periods, it is determined to be a control transient disturbance. Only when the residual appears continuously for ≥2 sampling periods in at least two modalities, has the same sign of change, or corresponds to the same state feature change, is it retained as valid latent variable data.

[0052] Optionally, the latent variable back estimation includes: Sensitive residual data are obtained by performing residual sensitivity matching based on latent variable data; In one embodiment, the system filters predefined residual terms corresponding to the latent variables from the multimodal residual data based on the latent variable type in the latent variable data. If the latent variable type is friction coefficient anomaly, then slip displacement residual, tangential force residual, pressure center offset residual, and contact stability residual are extracted; if it is elastic parameter anomaly, then deformation residual, contact area residual, and pressure distribution diffusion residual are extracted; if it is center of gravity offset, then torque residual, rotation angle residual, and transport posture residual are extracted; if it is internal structure anomaly, then local pressure mutation residual, discontinuous deformation residual, and force response mutation residual are extracted. The system establishes a mapping relationship between latent variable types and corresponding sensitive residual terms to obtain sensitive residual term data.

[0053] Parameter inversion is performed based on the sensitive residual data to obtain the latent variable parameter data; In one embodiment, the system selects the corresponding inversion method based on the physical type of the sensitive residual. For slip-type sensitive residuals, the friction coefficient range is inferred from the tangential force, normal force, and slip trend (e.g., the upper limit of the friction coefficient is determined based on the ratio of tangential force to normal force, and the range is narrowed by combining the slip duration), where the slip trend is an increase in slip displacement over 2-3 consecutive sampling periods. For deformation-type sensitive residuals, the elastic parameters are inferred from the applied force, changes in contact area, and deformation. For attitude and torque-type sensitive residuals, the direction and distance of the center of gravity offset are inferred from the torque direction, changes in rotation angle, and the position of the gripping point. For local abrupt change-type sensitive residuals, the abnormal areas of the internal structure are inferred from the location of pressure abrupt change, the region of deformation discontinuity, and the moment of force response abrupt change. The system obtains latent variable parameter data, including at least the parameter type, parameter range, inversion basis, and corresponding residual terms.

[0054] In one embodiment, the system performs parameter inversion based on sensitive residual data. Taking friction characteristic inversion as an example, when the tangential force residual is continuously greater than a preset threshold (e.g., greater than 1.5 times the historical average), and the slip displacement increases within three consecutive sampling periods while the contact pressure remains stable, the system initially determines the upper limit of the friction coefficient based on the current ratio range of tangential force to normal force. For example, when the tangential force is 20N and the normal force is 50N, the corresponding upper limit of the friction coefficient is 0.4. Combined with the slip duration (e.g., continuously for 3-5 sampling periods), the friction coefficient range is narrowed to 0.2-0.4.

[0055] Taking the inversion of material elastic parameters as an example, when the contact area increases by more than 10% in three consecutive sampling periods and the deformation continues to increase (for example, from 2 mm to 5 mm), the system infers the range of elastic parameters based on the relationship between the applied force and the deformation. For example, when the applied force is 30 N, the corresponding deformation is 5 mm, and the range of elastic parameters can be determined to be 6-15 N / mm.

[0056] Taking the center of gravity shift inversion as an example, when the torque residual exceeds a preset threshold (e.g., more than 1.5 times the historical average) within two consecutive sampling periods, and the rotation angle deviation continues to increase (e.g., from 3° to 8°), the system determines the center of gravity shift direction based on the torque direction and the shift distance range based on the rotation amplitude, for example, a shift range of 5mm-15mm. The system encapsulates the obtained parameter range, corresponding residual terms, and inversion basis into latent variable parameter data.

[0057] The residual explanatory power is calculated based on the latent variable parameter data to obtain the residual explanatory power data; In one embodiment, the system substitutes the latent variable parameter data into the aforementioned action correlation relationship, and re-deduces the target response under the same action command, initial pose, and contact state to obtain residual data after introducing the latent variable parameters. The system compares this residual data with the original multimodal residual data and calculates the change ratio of the amplitude of each residual term. The residual reduction ratio is the decrease ratio of the amplitude of the same residual term before and after introducing the latent variable parameters; the number of explained residual terms is the number of residual terms whose residual amplitude gradually decreases over 2-3 consecutive sampling periods; and the newly added residuals are residual terms generated after introducing the latent variable parameters and whose amplitude exceeds a preset threshold. The system obtains residual interpretability data based on the above results.

[0058] Based on the residual explanatory power data, parameter convergence screening was performed to obtain the latent variable estimation data.

[0059] In one embodiment, the system sorts candidate latent variable parameters according to their residual explanatory power (e.g., reference residual reduction ratio) and eliminates parameters that are physically infeasible or cannot simultaneously reduce the amplitude of multiple residual terms. The candidate latent variable parameters are parameter values ​​or ranges obtained through parameter inversion based on the latent variable type and corresponding residual combination in the latent variable data. For example, the friction coefficient parameter must not exceed the preset physical parameter range for material contact (e.g., friction coefficient between 0 and 1), the elastic parameter must not deviate from the actual deformation recovery direction or trend, the center of gravity offset direction should have the same sign as the torque residual direction or an angle less than a preset threshold, and abnormal internal structural regions should correspond to the location of pressure abrupt changes. If, after introduction, multiple candidate latent variable parameters can reduce the amplitude of the corresponding residual term within 2-3 consecutive sampling periods, and the number of explained residual terms reaches a preset threshold (e.g., not less than the dominant mode residual term), while the number of newly added residuals does not exceed the preset threshold, then it is determined that all multiple candidate latent variable parameters can explain the residuals. If multiple candidate latent variable parameters can explain the residuals, the parameter with more residual modal coverage, fewer conflicting residuals, and which satisfies the residual descent condition across multiple action stages is selected first. The system encapsulates the selected latent variable type, parameter value or parameter range, residual explained data, and corresponding residual source into latent variable estimation data.

[0060] Optionally, the parameter inversion includes: The inversion method is selected based on the sensitive residual data, and the inversion method data is obtained; In one embodiment, the system first identifies the physical properties and source modes of sensitive residual terms. If the sensitive residual term includes slip displacement residual, tangential force residual, and pressure center offset residual, friction parameter inversion is selected; if it includes deformation residual, contact area expansion residual, and pressure distribution diffusion residual, elastic parameter inversion is selected; if it includes torque residual, rotation angle residual, and transport posture residual, center of gravity parameter inversion is selected; if it includes local pressure abrupt change residual, discontinuous deformation residual, and force response abrupt change residual, structural parameter inversion is selected. If the same residual term satisfies multiple conditions simultaneously, multiple inversion methods are retained.

[0061] The sensitive residual data are inverted based on the inversion method data to obtain the latent variable parameter data. The inversion process includes at least one of friction parameter inversion, elastic parameter inversion, center of gravity parameter inversion, and structural parameter inversion.

[0062] In one embodiment, when the inversion method is friction parameter inversion, the system extracts the normal force, tangential force, slip displacement, and slip duration. If, during the clamping or moving stage, the actual tangential bearing capacity is lower than the predicted value and the target object undergoes continuous displacement along the contact surface direction, a low friction trend is determined. The system inversely calculates the friction coefficient range based on the critical relationship between the tangential force and the normal force, for example, using... As the basis for critical estimation, The coefficient of friction, For tangential force, The normal force is used, and the interval is narrowed by combining the slip velocity and duration to obtain friction parameter data.

[0063] When the inversion method is elastic parameter inversion, the system extracts the applied normal force, visual deformation, tactile contact area, and deformation recovery. If, under the same clamping force, the actual deformation is greater than the predicted deformation, and the contact area increases by more than a preset proportion (e.g., more than 10%) within 2-3 consecutive sampling periods, the target object is determined to have low stiffness. The system infers the elastic parameters based on the response relationship between force and deformation, for example, using... As the basis for elasticity estimation, For equivalent elastic parameters, In order to apply force, The deformation is used as the variable, and the deformation recovery ratio during the release phase is combined to determine whether there is elastic hysteresis in the material, thus obtaining elastic parameter data.

[0064] When the inversion method is center of gravity parameter inversion, the system extracts the torque residual, rotation angle residual, gripping point position, and changes in transport posture. If, under symmetrical gripping or expected linear transport conditions, the target object exhibits a continuous rotational deviation, and the torque direction is consistent with the rotation direction (i.e., the direction signs are the same or the included angle is less than a preset threshold (e.g., 15°), then a center of gravity offset is determined to exist. The system determines the direction of the center of gravity offset based on the torque direction and estimates the degree of offset based on the rate of change of the rotation angle, for example, using... As the basis for offset estimation, The offset distance of the center of gravity relative to the grasping point. For torque, The force is used to obtain the center of gravity parameter data.

[0065] When the inversion method is structural parameter inversion, the system extracts the locations of local pressure abrupt changes, visual discontinuous deformation regions, and force feedback abrupt change times from the tactile array. The abrupt change is defined as the change in pressure between adjacent sampling times exceeding a preset threshold (e.g., 1.5 times the historical average). If the change in local pressure between adjacent sampling times exceeds the preset threshold (e.g., 1.5 times the historical average), and this location coincides with a visual deformation discontinuity region, it is determined that an internal cavity, weak layer, or heterogeneous structure exists. The system determines the location of structural anomalies based on the pressure abrupt change region, determines the anomaly intensity based on the abrupt change amplitude and the degree of deformation discontinuity, and determines whether the anomaly was triggered during contact loading based on the force response abrupt change time, thus obtaining structural parameter data. The method for determining the abnormal intensity includes the system first calculating the pressure difference between adjacent tactile units within the abnormal area. If the pressure difference is greater than 1.5 times the historical average pressure, a local pressure mutation is determined. The system then calculates the deformation difference between the abnormal area and adjacent areas. If the deformation difference is greater than a preset deformation threshold, for example, greater than 20% of the average deformation of adjacent areas, a deformation discontinuity is determined. The system classifies the pressure mutation amplitude and the degree of deformation discontinuity into three levels: low, medium, and high. When both are low, the structural abnormal intensity is determined to be low; when either is high or both are medium, the structural abnormal intensity is determined to be medium; and when both are high, the structural abnormal intensity is determined to be high. The method of determining whether an anomaly is triggered during contact loading based on the moment of force response change includes: the system divides the action process into an approach phase, a contact loading phase, a stable clamping phase, and a moving phase, and takes the moment when the tactile contact unit first exceeds the contact threshold as the contact loading start point; the system acquires the moment of force response change in the force feedback data. If the change in normal force, tangential force, or torque exceeds a preset change threshold during the contact loading phase, for example, greater than 1.5 times the historical average, and the time difference between the moment of change and the moment of local pressure change or deformation discontinuity is less than a preset window, for example, less than 2 sampling periods, then it is determined that the structural anomaly is triggered during contact loading. If the force response change occurs before contact or after contact loading and has no time correspondence with local pressure or deformation anomalies, then it is not determined to be a contact loading triggered anomaly.

[0066] The system organizes friction parameter data, elastic parameter data, center of gravity parameter data, and structural parameter data according to their corresponding residual sources. If multiple parameters correspond to the same residual combination, they are retained in parallel, and their corresponding residual combinations, the stage of occurrence, and the inversion method are recorded. If different parameters correspond to different residual sources, for example, one parameter corresponds only to slip-related residuals while another corresponds to stress-moment-related residuals, their correspondence is recorded separately. Parameters that cannot simultaneously correspond to multiple residual terms are marked as locally interpreted parameters. The system encapsulates the above parameters and their correspondences into latent variable parameter data, which includes at least the latent variable type, parameter estimate or parameter range, inversion method, and corresponding sensitive residual terms.

[0067] Optionally, the calculation of the residual explanation degree includes: Based on the latent variable parameter data, the motion operation is reconstructed to obtain motion reconstruction data; In one embodiment, the system reintroduces latent variable parameter data into the original motion prediction process. While maintaining the original motion command, initial pose, gripping point, contact phase, and execution time window unchanged, the system reconstructs the state response of the target object. For example, the friction coefficient parameter is substituted into the slip response relationship, the elastic parameter into the deformation response relationship, the center of gravity offset parameter into the torque response relationship, and the structural anomaly parameter into the local pressure response relationship. These response relationships are state change mapping relationships established based on the aforementioned motion correlations. The system then re-deduces the target object's pose change, contact pressure distribution, force feedback change, deformation, and motion trajectory under this motion, obtaining motion reconstruction data.

[0068] Residual calculations are performed based on action reconstruction data and execution feedback data to obtain latent variable residual data; In one embodiment, the system compares the motion reconstruction data with the actual execution feedback data within the same motion stage and the same time window, and calculates the residual after introducing latent variable parameters. If the motion reconstruction data includes reconstructed pose, reconstructed center of pressure, reconstructed force feedback, and reconstructed trajectory, then the differences are calculated with the actual pose, actual center of pressure, actual force feedback, and actual trajectory, respectively. This can be expressed as: ,in, The residuals after introducing latent variables, This is the actual feedback status. The state is reconstructed for the action. The system obtains the latent variable residual data.

[0069] Based on the latent variable residual data and the latent variable data, residual reduction judgment is performed to obtain the residual reduction data; In one embodiment, the system compares the latent variable residual data with the original multimodal residual data to determine whether the residuals decrease after introducing the latent variable parameter. For sensitive residual terms corresponding to the latent variable data, if the magnitude of the latent variable residual is less than the magnitude of the original residual, and the decreasing trend remains consistent within a continuous time window, then it is determined that the latent variable parameter can explain the corresponding residual. This can be expressed as: ,in For the original residual, For latent variable residuals; when A value greater than 0 indicates that the residual has been reduced. The system outputs the magnitude, percentage, and duration of the residual reduction.

[0070] Residual term coverage is performed based on the residual reduction data to obtain residual coverage data; In one embodiment, the system statistically analyzes the residual terms and their corresponding modes that can be effectively reduced by the same latent variable parameter. If the friction parameter can simultaneously reduce the slip displacement residual, tangential force residual, and pressure center offset residual, it is determined that it covers multiple modes including visual, force feedback, and tactile. If the elastic parameter can reduce the deformation residual, contact area residual, and pressure distribution diffusion residual, it is determined that it covers the deformation and tactile response related residuals. The system records the effectively reduced residual terms, corresponding modes, occurrence stages, and coverage duration to obtain residual coverage data, which is used to characterize the interpretability range of the latent variable parameter.

[0071] The residual reduced data and residual covered data are used to evaluate the explanatory power of the latent variable parameter data, thus obtaining the residual explanatory power data.

[0072] In one embodiment, the system evaluates the interpretability of latent variable parameters based on the residual reduction magnitude, reduction ratio, number of covered residual terms, number of covered modes, and coverage duration. If a latent variable parameter can continuously reduce multiple modal residuals within multiple time windows, and the covered residual terms are consistent with the sensitive residual terms of that latent variable type, it is determined to have high interpretability; if it only reduces a single residual term or the reduction is unstable, it is determined to have low interpretability. The system obtains residual interpretability data, which includes at least the latent variable type, parameter value or parameter range, residual reduction ratio, covered residual terms, covered modes, interpretability level, and interpretability label.

[0073] Optionally, the latent variable back estimation includes: Spatial mapping is performed on the latent variable data to obtain spatially mapped data; In one embodiment, the system uniformly maps the latent variable type, residual source mode, occurrence stage, residual amplitude, residual duration, and confidence level from the latent variable data to the latent variable state space. This latent variable state space can be divided into different subspaces based on frictional characteristics, elastic characteristics, center of gravity distribution, and structural anomalies. For example, low-friction latent variables / friction coefficient latent variables are mapped to the friction coefficient subspace, soft deformation latent variables / elastic modulus latent variables are mapped to the elastic parameter subspace, and center of gravity offset latent variables are mapped to the mass distribution subspace. The system represents each latent variable as a state vector Z=[c,r,s,t,p], where c is the latent variable type, r is the associated residual term, s is the occurrence stage, t is the duration, and p is the confidence level, thus obtaining the spatial mapping data. The system determines confidence based on the proportion of valid sampling points of the residuals corresponding to latent variables within a preset observation window. For example, if the observation window is set to 5 sampling periods, the system counts the number of sampling points that meet the residual validity criteria within this time window. The residual validity criteria are that the residual amplitude is greater than a preset threshold and the residuals of at least two modes change in the same direction. Let n be the number of sampling points that meet the above conditions. Then, the confidence level p is determined as the ratio of n to the total number of sampling points in the observation window, i.e., p = n / 5, where p ranges from 0 to 1.

[0074] The spatial mapping data is reconstructed using a probability distribution to obtain the reconstructed distribution data. In one embodiment, the system divides the latent variable parameter space into several parameter intervals and maps the state vectors in the spatial mapping data to the corresponding intervals according to the latent variable type. For each interval, the reduction and confidence level of the corresponding residual terms are statistically analyzed. When an interval reduces the residual amplitude and has a high confidence level within 2-3 consecutive sampling periods, the weight of that interval is increased; if the residual does not decrease or new residuals appear, its weight is decreased. The system normalizes the weights of each interval so that the sum of the weights is 1, thereby obtaining the probability distribution of the latent variable parameters in each interval, forming the distributed reconstruction data.

[0075] Perform correlation reasoning on the distributed reconstructed data to obtain correlation reasoning data; In one embodiment, parameter intervals with probability values ​​greater than a preset threshold (e.g., greater than 0.2) are selected as candidate intervals, and correlation analysis is performed in conjunction with the corresponding latent variable types and their residual sources. If multiple parameter intervals correspond to the same residual combination and exhibit consistent probability change trends within the same action phase, an association relationship is established. Consistent probability change trends mean that within 2-3 consecutive time windows, the probability values ​​of each parameter interval change in the same direction, i.e., all increasing or all decreasing, and the difference in probability change between adjacent time windows is less than a preset threshold (e.g., less than 0.05). If different latent variable parameters coexist within adjacent time windows and their corresponding residual terms have cross-influence, they are identified as associated latent variables. Cross-influence means that the residual terms corresponding to different latent variable parameters change simultaneously within the same time window, and the change in at least one latent variable parameter can cause a change in the amplitude of another residual term. For example, when a change in friction parameters affects not only the slip residual but also the torque residual, or a change in elastic parameters affects both the deformation residual and the contact area residual, and these effects persist for 2-3 consecutive sampling periods, cross-influence is identified. The system generates the association structure between latent variables based on the association relationship, obtaining associated inference data.

[0076] By performing a reverse iteration of change optimization based on the correlation inference data, we obtain the latent variable estimation data.

[0077] In one embodiment, the system evaluates the probability value, the number of corresponding residual terms, and the number of associated latent variables for each latent variable parameter interval. It prioritizes parameter intervals with probability values ​​greater than a preset threshold (e.g., greater than 0.3) and a large number of corresponding residual terms as candidate iteration intervals. When multiple intervals meet the conditions, the interval that is associated with other latent variables—that is, the parameter interval that is associated with at least one other latent variable in the association inference data—is selected as the priority iteration interval. If multiple candidate intervals still exist, the parameter interval whose probability value continuously increases over 2-3 consecutive time windows is prioritized. Through the above, the latent variable parameter intervals that need to be prioritized for iteration are determined. The system substitutes the parameter interval into the aforementioned state deduction process and recalculates the corresponding residuals. If the residual amplitude does not decrease, the adjacent interval in the parameter space is selected as the latent variable parameter interval to be iterated first, and this interval is sequentially substituted into the state deduction process to calculate the corresponding residuals. The interval that causes the residual amplitude to decrease within 2-3 consecutive sampling periods is selected as the new current parameter interval. If the residual decrease condition is not met in either direction, the search range is expanded to the next adjacent interval. If the residual amplitude decreases within 2-3 consecutive sampling periods without generating new residuals, the parameter interval is retained. The system repeats the above process until the parameter interval change is less than the preset step size or the maximum number of iterations is reached, thus obtaining the latent variable estimation data.

[0078] Optionally, the action decision correction includes: S41. Perform decision analysis on the latent variable estimation data to obtain decision analysis data; In one embodiment, the system reads the latent variable type, parameter range, confidence level, and corresponding residual source from the latent variable estimation data, and converts them into action decision constraints. If the latent variable is the friction coefficient, it is interpreted as the contact stability index being lower than a preset threshold, indicating an increased risk of slippage. If the latent variable is the elastic modulus, it is interpreted as an increased risk of pressure loss and a decreased upper limit of clamping force. If the latent variable is the center of gravity shift, it is interpreted as the distance between the gripping point and the center of gravity shift exceeding a preset threshold, indicating a tendency for the moving posture to become unstable. If the latent variable is an internal structural anomaly, it is interpreted as an increased risk of localized force concentration. The system obtains decision analysis data, which includes the risk type, the stage of the action affecting it, the parameters of the action being affected, and the correction direction.

[0079] S42. Based on the decision impact data, conduct a reassessment of action risk to obtain action risk data; In one embodiment, the system extracts relevant state variables for judgment during the corresponding action phase: when the ratio of tangential force to normal force exceeds a preset threshold and the slip displacement residual increases within 2-3 consecutive sampling periods, slip risk is determined; when the deformation exceeds a preset deformation threshold, pressure loss risk is determined; when the distance between the gripping point and the center of gravity exceeds a preset distance threshold, attitude instability risk is determined; when the contact area coincides with the pressure change area, structural anomaly risk is determined. The system determines the corresponding risk level based on the magnitude of each risk indicator exceeding the threshold using preset grading thresholds, thus obtaining action risk data.

[0080] S43. Determine the correction target based on the action risk data to obtain action correction target data; In one embodiment, if the slippage risk level exceeds a preset threshold, improving contact stability is used as the correction target; if the pressure loss risk level exceeds a preset threshold, reducing local contact pressure is used as the correction target; if the attitude instability risk level exceeds a preset threshold, reducing torque or adjusting gripping balance is used as the correction target; if the structural anomaly risk level exceeds a preset threshold, avoiding pressure abrupt change areas is used as the correction target. When multiple risks exist, the system prioritizes the correction target corresponding to the risk with the highest risk level and uses the targets corresponding to other risks as constraints to obtain action correction target data.

[0081] S44. Adjust the motion parameters based on the motion correction target data to obtain motion correction data; In one embodiment, the system adjusts the corresponding motion parameters based on the motion correction target data. If the correction target is to improve contact stability, the clamping force is increased during the clamping phase to not exceed a preset force control upper limit, and the transfer acceleration is reduced. If the correction target is to reduce local contact pressure, the clamping force is reduced or the contact area is expanded so that the contact pressure is lower than a preset pressure threshold. If the correction target is to reduce torque or adjust the gripping balance, the gripping point is adjusted towards the estimated center of gravity so that the offset distance is less than a preset threshold. If the correction target is to avoid areas of sudden pressure changes, the contact area is reselected or the path is adjusted so that the contact position does not coincide with the area of ​​sudden pressure changes. The system obtains motion correction data.

[0082] S45. Update the action decision based on the action correction data to obtain action decision correction data.

[0083] In one embodiment, the system writes the corrected parameters such as gripping force, grasping point, contact angle, and movement speed into the current motion control command, and simultaneously updates the execution path and constraints of the corresponding motion stage. The updated motion is then validated for feasibility. If the corrected parameters meet the robotic arm's range of motion, force control upper limit, and obstacle avoidance constraints, the update is confirmed as effective. If any parameters do not meet the constraints, they are rolled back or readjusted. The system generates motion decision correction data, which includes the updated motion parameters, execution path, and constraints.

[0084] Optionally, this application also provides an embodied intelligent multimodal perception and decision-making system for executing the embodied intelligent multimodal perception and decision-making method described above, the embodied intelligent multimodal perception and decision-making system comprising: The action prediction module is used to acquire multimodal data of embodied interaction and perform action prediction based on the multimodal data of embodied interaction to obtain action prediction data; The multimodal residual calculation module is used to acquire execution feedback data and calculate state deviation based on action prediction data and execution feedback data to obtain multimodal residual data. The latent variable estimation module is used to identify residuals in multimodal residual data to obtain latent variable data; and to perform latent variable back-estimation based on the latent variable data to obtain latent variable estimated data. The action decision correction module is used to correct action decisions based on latent variable estimation data, and obtain action decision correction data.

[0085] Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended application documents rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of the equivalents of the application documents be incorporated into the invention.

[0086] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. An embodied intelligence multimodal perception and decision-making method, characterized in that, Includes the following steps: Acquire multimodal data of embodied interaction, and perform action prediction based on the multimodal data of embodied interaction to obtain action prediction data; Obtain execution feedback data, and calculate state deviation based on action prediction data and execution feedback data to obtain multimodal residual data; Residual identification is performed on multimodal residual data to obtain latent variable data; Latent variable estimation is performed based on latent variable data to obtain latent variable estimation data; Action decision correction data is obtained by adjusting the action decision based on the latent variable estimation data.

2. The embodied intelligence multimodal perception and decision-making method according to claim 1, characterized in that, The action prediction includes: Cross-modal state feature extraction is performed based on embodied interaction multimodal data to obtain interaction state feature data; Action association modeling is performed on the interaction state feature data to obtain action association data; Based on the action-related data, the expected state is deduced to obtain the expected state data; Multimodal consistency verification is performed based on the expected state data to obtain action prediction data.

3. The embodied intelligence multimodal perception and decision-making method according to claim 1, characterized in that, The state deviation calculation includes: The action prediction data and execution feedback data are matched at the feedback time to obtain the time-matching data. State deviation is calculated based on time-matching data to obtain state deviation data, which includes visual state deviation calculation, tactile state deviation calculation, force feedback deviation calculation, and motion trajectory deviation calculation. The state deviation data is processed by residual coupling to obtain multimodal residual data.

4. The embodied intelligence multimodal perception and decision-making method according to claim 1, characterized in that, The residual identification includes: Residual mode determination is performed on the multimodal residual data to obtain residual mode data; Based on the residual modal data, cross-modal residual combination is performed to obtain residual combination data; Perform row latent variable type matching on the residual combination data to obtain latent variable type data; The latent variable data is obtained by eliminating the source of the latent variable based on its type.

5. The embodied intelligence multimodal perception and decision-making method according to claim 1, characterized in that, The back-estimation of latent variables includes: Sensitive residual data are obtained by performing residual sensitivity matching based on latent variable data; Parameter inversion is performed based on the sensitive residual data to obtain the latent variable parameter data; The residual explanatory power is calculated based on the latent variable parameter data to obtain the residual explanatory power data; Based on the residual explanatory power data, parameter convergence screening was performed to obtain the latent variable estimation data.

6. The embodied intelligence multimodal perception and decision-making method according to claim 5, characterized in that, The parameter inversion includes: The inversion method is selected based on the sensitive residual data, and the inversion method data is obtained; The sensitive residual data are inverted based on the inversion method data to obtain the latent variable parameter data. The inversion process includes at least one of friction parameter inversion, elastic parameter inversion, center of gravity parameter inversion, and structural parameter inversion.

7. The embodied intelligence multimodal perception and decision-making method according to claim 5, characterized in that, The calculation of the residual explanation degree includes: Based on the latent variable parameter data, the motion operation is reconstructed to obtain motion reconstruction data; Residual calculations are performed based on action reconstruction data and execution feedback data to obtain latent variable residual data; Based on the latent variable residual data and the latent variable data, residual reduction judgment is performed to obtain the residual reduction data; Residual term coverage is performed based on the residual reduction data to obtain residual coverage data; The residual reduced data and residual covered data are used to evaluate the explanatory power of the latent variable parameter data, thus obtaining the residual explanatory power data.

8. The embodied intelligence multimodal perception and decision-making method according to claim 1, characterized in that, The back-estimation of latent variables includes: Spatial mapping is performed on the latent variable data to obtain spatially mapped data; The spatial mapping data is reconstructed using a probability distribution to obtain the reconstructed distribution data. Perform correlation reasoning on the distributed reconstructed data to obtain correlation reasoning data; By performing a reverse iteration of change optimization based on the correlation inference data, we obtain the latent variable estimation data.

9. The embodied intelligence multimodal perception and decision-making method according to claim 1, characterized in that, The action decision correction includes: Decision analysis is performed on the latent variable estimation data to obtain decision analysis data; Based on the decision impact data, the action risk is reassessed to obtain action risk data; Based on the action risk data, the correction target is determined, and the action correction target data is obtained; The motion parameters are adjusted based on the motion correction target data to obtain the motion correction data; The action decision is updated based on the action correction data to obtain action decision correction data.

10. An embodied intelligent multimodal perception and decision-making system, characterized in that, For executing the embodied intelligent multimodal perception and decision-making method as described in claim 1, the embodied intelligent multimodal perception and decision-making system includes: The action prediction module is used to acquire multimodal data of embodied interaction and perform action prediction based on the multimodal data of embodied interaction to obtain action prediction data; The multimodal residual calculation module is used to acquire execution feedback data and calculate state deviation based on action prediction data and execution feedback data to obtain multimodal residual data. The latent variable estimation module is used to identify residuals in multimodal residual data to obtain latent variable data; and to perform latent variable back-estimation based on the latent variable data to obtain latent variable estimated data. The action decision correction module is used to correct action decisions based on latent variable estimation data, and obtain action decision correction data.