Model training method, robot and computer readable storage medium

CN122807875APending Publication Date: 2026-09-25SHENZHEN LINGSI ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610959372.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,在视觉观测数据发生变化时,模型仅凭视觉观测数据无法区分变化是由物体的真实运动还是视角变化导致,从而导致模型对空间关系进行错误推导,进而导致模型输出的预测动作偏差较大

Benefits of technology

[0008]本申请提供一种基于观测部件定位的模型训练方法,所述方法基于第一历史动作数据、第一位姿数据、第二历史动作数据、第二位姿数据、目标相对位姿数据以及历史视觉观测数据,对模仿学习模型进行基于观测部件定位的模型训练;其中,所述第一历史动作数据与所述第二历史动作数据分别为执行端在第一时刻与下一时刻的历史执行动作数据;所述第一位姿数据与所述第二位姿数据分别为执行端与视觉观测部件在所述第一时刻的位姿数据;所述目标相对位姿数据根据所述第一位姿数据与所述第二位姿数据计算,用于表征所述视觉观测部件与所述执行端在所述第一时刻的视角空间关系的相对位姿;所述历史视觉观测数据由所述视觉观测部件采集。通过上述方式,本申请采集视觉观测部件的第二位姿数据,并根据第二位姿数据与执行端的第一位姿数据,计算出执行端与视觉观测部件的相对位姿数据。根据相对位姿数据,确定执行端在视觉观测部件坐标系中的姿态。由于相对位姿仅随执行端的真实运动而变化,而不随视角变化而变化,使得模型基于相对位姿确定执行端与观测部件的真实空间关系,并基于真实空间关系与执行端的动作之间的映射关系进行模型训练,从而避免因视觉变化对空间关系进行错误推导,进而降低了模型的动作预测偏差,提高了模型的动作推理精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807875A_ABST
    Figure CN122807875A_ABST
Patent Text Reader

Abstract

The application provides a model training method, a robot and a computer readable storage medium. The method collects second pose data of a visual observation component, and calculates relative pose data of an execution end and the visual observation component according to the second pose data and first pose data of the execution end. According to the relative pose data, the attitude of the execution end in the coordinate system of the visual observation component is determined. Since the relative pose only changes with the real movement of the execution end and does not change with the visual angle, the model determines the real space relationship between the execution end and the observation component based on the relative pose, and performs model training based on the mapping relationship between the real space relationship and the action of the execution end, thereby avoiding incorrect deduction of the space relationship due to visual changes, reducing the action prediction deviation of the model, and improving the action reasoning accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more particularly to a model training method, a robot, and a computer-readable storage medium. Background Technology

[0002] Currently, in the training process of robot imitation learning models, to enable robots to more realistically simulate user operations, visual observation data (such as RGB images and depth data) is collected from the execution end (such as a handheld gripper) during task completion by using a perspective observation device (such as a head-mounted camera) worn by the operator to mimic a first-person perspective and human observation. Then, the visual observation data and the execution end's state information (such as pose and degree of opening / closing) are used as input features and fed into the robot imitation learning model, allowing the model to infer the sequence of actions related to the execution end's task completion based on the observation data and state information.

[0003] However, when visual observation data changes, the model cannot distinguish whether the change is due to the actual movement of the object or a change in perspective, leading to incorrect inferences about spatial relationships and consequently, significant deviations in the model's predicted actions. For example, if the wearer turns their head to the left, and the corresponding image shows the gripper moving to the right, the model might misinterpret this as the gripper moving to the right. Summary of the Invention

[0004] The main purpose of this application is to provide a model training method, a robot, and a computer-readable storage medium, which aims to reduce the action prediction bias of a robot's imitation learning model.

[0005] To achieve the above objectives, this application provides a model training method, which includes the following steps: Based on the first historical action data, the first pose data, the second historical action data, the second pose data, the target relative pose data, and the historical visual observation data, the imitation learning model is trained based on the localization of the observed parts. Wherein, the first historical action data and the second historical action data are respectively the historical execution action data of the execution end at the first moment and the next moment; the first pose data and the second pose data are respectively the pose data of the execution end and the visual observation component at the first moment; the target relative pose data is calculated based on the first pose data and the second pose data, and is used to characterize the relative pose of the visual observation component and the execution end in the viewpoint space relationship at the first moment; the historical visual observation data is collected by the visual observation component.

[0006] In addition, to achieve the above objectives, this application also provides a robot, comprising: The visual observation component is used to collect historical visual observation data of the execution end when performing historical actions; The first positioning module is used to locate the first pose data of the execution end; The second positioning module is disposed on the visual observation component and is used to locate the second pose data of the visual observation component. The processor, the memory, and the action generation program based on observation component localization stored in the memory and executable by the processor, wherein when the action generation program based on observation component localization is executed by the processor, the steps of the model training method based on observation component localization as described above are implemented.

[0007] In addition, to achieve the above objectives, this application also provides a computer-readable storage medium storing an action generation program based on observation component localization, wherein when the action generation program based on observation component localization is executed by a processor, it implements the steps of the model training method based on observation component localization as described above.

[0008] This application provides a model training method based on observation component localization. The method trains an imitation learning model based on observation component localization using first historical action data, first pose data, second historical action data, second pose data, target relative pose data, and historical visual observation data. The first and second historical action data represent the historical execution action data of the execution end at a first and next time moment, respectively. The first and second pose data represent the pose data of the execution end and the visual observation component at the first time moment, respectively. The target relative pose data is calculated based on the first and second pose data to characterize the relative pose of the visual observation component and the execution end at the first time moment. The historical visual observation data is collected by the visual observation component. Through this method, this application collects the second pose data of the visual observation component and calculates the relative pose data between the execution end and the visual observation component based on the second pose data and the first pose data of the execution end. Based on the relative pose data, the posture of the execution end in the coordinate system of the visual observation component is determined. Since the relative pose changes only with the actual movement of the actuator and not with the viewpoint, the model determines the real spatial relationship between the actuator and the observation component based on the relative pose, and trains the model based on the mapping relationship between the real spatial relationship and the action of the actuator. This avoids incorrect inference of spatial relationship due to visual changes, thereby reducing the model's action prediction bias and improving the model's action inference accuracy. Attached Figure Description

[0009] Figure 1 A flowchart illustrating a model training method provided in this application; Figure 2 A flowchart illustrating another model training method provided in this application; Figure 3 A flowchart illustrating another model training method provided in this application; Figure 4 A flowchart illustrating the head-mounted visual observation component provided in this application; Figure 5 A schematic diagram of a robot module provided in this application.

[0010] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0013] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0015] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0016] Reference Figure 1 , Figure 1 This is a flowchart illustrating a model training method provided in this application.

[0017] like Figure 1 As shown, the specific training method for this model includes: Step S101: Based on the first pose data of the execution end and the second pose data of the visual observation component at the first moment, the relative pose of the target is calculated. The relative pose of the target represents the spatial relationship of the viewpoint between the visual observation component and the execution end at the first moment.

[0018] In some embodiments, the execution end refers to the terminal device used to perform physical or simulated operations. In practical applications, this could be the end flange of an industrial robotic arm, the hand of a humanoid robot, or a gripper simulator held by the operator during the data acquisition phase. The visual observation component refers to the sensing unit used to acquire visual information about the environment, such as an RGB camera, depth camera, or multimodal sensor module mounted on the operator's or robot's head. The first moment refers to any sampling time point in the training data sequence.

[0019] Traditional visual imitation learning models struggle to accurately infer absolute scale and depth information from two-dimensional image pixels, leading to grasping failures or collisions. To address this issue, this embodiment calculates the relative pose data of the target between the visual observation component and the execution end, determining the viewpoint spatial relationship between them. This establishes a geometric association between the visual observation component and the execution end, providing a clear scale reference frame and spatial constraints for the imitation learning model. Therefore, by introducing the relative pose of the target between the visual observation component and the execution end into the training process of the imitation learning model, the model no longer relies solely on visual features but also incorporates the precise spatial transformation relationship between the perceptual observation viewpoint and the execution end's ontology. This solves the scale ambiguity problem under purely visual input and improves the imitation learning model's ability to understand the distance and size of objects from a first-person perspective.

[0020] Step S102: Based on the first pose data, the target relative pose data, the historical visual observation data, the first historical action data of the execution end at the first moment, and the second historical action data at the next moment, the imitation learning model is trained based on the localization of the observation component; wherein, the historical visual observation data is collected by the visual observation component.

[0021] In some embodiments, historical visual observation data includes image sequences, depth maps, or point cloud data acquired by the visual observation component at a first moment and previous moments. This historical visual observation data provides semantic information about the environment. The first pose data and the first historical action data together constitute the ontological state description of the actuator at the current moment. The first historical action data may include the opening and closing degree of the gripper, joint angles, or current movement speed. The second historical action data serves as an output label, representing the action the actuator should perform at the next moment, which can be a pose increment or joint torque several future moments. Further imitation learning model training is performed by combining the target relative pose data, i.e., predicting the future action of the actuator using the current relative pose of the visual observation component and the actuator during training. This enables the model to achieve conditioned reflex mapping; for example, when the head is at a specific viewpoint and the hand is at a relative position X, action Y should be performed.

[0022] Compared to traditional model training schemes that only use vision and ontology states, this embodiment reduces the difficulty of geometric reasoning in complex scenes by introducing the relative pose of the target, thereby improving the model training efficiency and the success rate of reasoning.

[0023] It is understood that the above embodiments only use a data stream at a single moment as an example for illustration. In actual training, time window data containing multiple consecutive moments can be further constructed to capture dynamic motion trends.

[0024] By incorporating the target relative pose between the visual observation component and the execution end into the training of the imitation learning model, this embodiment transforms the geometric relationships in image pixels into learnable features. This enables the imitation learning model to establish a mapping between head viewpoint, hand position, and object distance, thereby solving the problem of the lack of absolute scale reference in first-person vision in traditional imitation learning. By introducing spatial relationships, not only is the ambiguity in scale estimation in monocular or multi-view vision eliminated, but the imitation learning model can also adaptively adjust its understanding of the imitation learning scene based on real-time relative pose when faced with different heights, wearing positions, or robotic arm configurations. This improves the inference success rate of the imitation learning model in complex and unstructured environments.

[0025] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating another model training method provided in an embodiment of this application.

[0026] like Figure 2 As shown, step S101 specifically includes: Step S1011: When the first pose data and the second pose data belong to different coordinate systems, the first pose data and the second pose data are transformed to the same target coordinate system based on the preset calibration parameters. Step S1012: Based on the first pose data and the second pose data in the target coordinate system, calculate the pose transformation of the visual observation component relative to the execution end, as the target relative pose data.

[0027] In some embodiments, when the actuator and the visual observation component are provided with positioning information by different sensing systems, or when the positioning methods of the actuator and the visual observation component are different, the first pose data representing the actuator and the second pose data representing the visual observation component will be in different reference coordinate systems. For example, the first pose data of the actuator is calculated by the forward kinematics of the joint encoder of the robotic arm itself, and its reference is the coordinate system of the robotic arm base; while the second pose data of the visual observation component is located by the inertial measurement unit integrated into the head-mounted device or an external optical tracking system, and its reference is the head local coordinate system or the global tracking coordinate system. In order to truly reflect the spatial relationship between the observation viewpoint and the operating body, a mapping relationship between the two heterogeneous coordinate systems is established by introducing preset calibration parameters, and the values ​​under the two different references are calculated.

[0028] In this embodiment, the preset calibration parameter can be a homogeneous transformation matrix containing rotation and translation components, used to characterize the fixed installation deviation in space between the origin of the actuator coordinate system and the origin of the vision observation component coordinate system. Specifically, the preset calibration parameter can be obtained through a hand-eye calibration algorithm, i.e., by solving for the observation data of the calibration plate under multiple postures; it can also be obtained by direct measurement based on a 3D CAD assembly model; or it can be estimated online through excitation actions during robot operation. The preset calibration parameter describes the static geometric constraints between the two sensors, which can then be used for subsequent coordinate unification. After completing coordinate system one, the first pose data and the second pose data can be represented in the same target coordinate system to perform vector or matrix operations on the two pose data. The pose transformation data calculated based on the unified pose data is the target relative pose data.

[0029] In this embodiment, the coordinate reference of heterogeneous sensors is unified by preset calibration parameters, which ensures that the calculated target relative pose has physical consistency and accuracy under different hardware configurations or installation methods, and provides a standardized spatial relationship input for the imitation learning model.

[0030] It is understood that this embodiment takes the robotic arm base system and the head-mounted camera system as examples to explain how to satisfy the initial heterogeneity of the two and the conditions that can be converted to the same reference through calibration parameters. The coordinate system of the execution end can also be the moving chassis coordinate system or the world coordinate system, and the coordinate system of the visual observation component can also be the camera coordinate system mounted on a third-party bracket.

[0031] For example, the step of converting the first pose data and the second pose data to the same target coordinate system based on preset calibration parameters includes: The coordinate system of the visual observation component to which the second pose data belongs is used as the target coordinate system; Based on the transformation matrix corresponding to the first pose data and the second pose data, the first pose data is transformed from the execution end coordinate system to the visual observation component coordinate system.

[0032] In some embodiments, in order to facilitate the establishment of a vision-centered embodied perception mapping for the model, so that the trained model conforms to the cognitive logic of first-person perspective imitation learning, the coordinate system of the visual observation component is used as the target coordinate system, that is, the spatial reference benchmark centered on visual perception. The pose information of the execution end is represented as the position relative to the operator's eyes, so that the absolute geometric quantity of the physical world is transformed into a relative perceptual quantity that conforms to the laws of embodied cognition.

[0033] During the training of the imitation learning model, the perception of limb position is not based on absolute world coordinates or fixed robotic arm base coordinates, but rather on an egocentric coordinate system with the operator's eyes as the origin. For example, when a person reaches out to grab a cup on a table, the relative information is that the cup is 30 centimeters to the lower left of the center of vision, rather than the absolute information of the cup's distance from the room's origin (2m, 1m, 0.8m). Therefore, by setting the visual observation component coordinate system as the target coordinate system, the first pose data input to the model represents the position and posture of the actuator relative to the operator's eyes (i.e., the visual observation component representation). Through this data representation method, the model is aligned with the visual semantics of the first-person perspective and human cognitive habits, thereby establishing a semantic association between the model's input features and the visual observation data. For example, when the operator's head rotates or the chassis moves, causing overall displacement, the spatial relationship of the actuator relative to the camera remains unchanged (i.e., the grasping intention remains unchanged), but the coordinate values ​​in the actuator's base system or world system change.

[0034] To avoid the model expending significant computational resources learning how to filter interference from head movements, and to reduce the difficulty of model fitting and training efficiency, this embodiment uses the visual observation component coordinate system as a reference and trains the model based on the mapping relationship of visual actions, thus eliminating the influence of head movements on the description of the relative relationship between the head and hands.

[0035] Specifically, the above transformation process includes the inverse operation and multiplication operation of the homogeneous transformation matrix: The second pose data represents the pose transformation matrix of the visual observation component in the world coordinate system. The first pose data represents the pose transformation matrix of the execution end in the world coordinate system. In this case, firstly Inverse is obtained Then combine it with Multiply to obtain the pose transformation matrix of the execution end in the coordinate system of the vision observation component. This matrix represents the final target relative pose data used for training.

[0036] In a specific embodiment, the transformation strategy centered on the observation perspective, in addition to using the coordinate system of the visual observation component itself as a reference in this embodiment, can also include an auxiliary link coordinate system that is rigidly connected to the visual observation component and parallel to the optical axis, and use it as the target coordinate system.

[0037] Understandably, this transformation strategy applies not only to data preprocessing during the training phase, but also to the model inference and deployment phase.

[0038] As one implementation, the method further includes: When the execution end is a single gripper, the single gripper pose information of the execution end at the first moment is collected as the first pose data; The opening and closing degree of the single gripper at the first moment of the execution end is collected as the first historical action data.

[0039] In some embodiments, a single gripper refers to a configuration containing only one independent, controllable end effector. Its specific form is not limited to common two-finger parallel grippers or three-finger adaptive grippers; it can also be any device with a single gripping or contact function, such as a vacuum suction cup, electromagnetic adsorption head, soft gripper, or single-finger probe. The single gripper pose information describes the geometric state of the single actuator in three-dimensional space, such as a six-degree-of-freedom vector containing three translational components (x, y, z) and three rotational components (roll, pitch, yaw), or expressed as a 4×4 homogeneous transformation matrix. Specifically, the single gripper pose information can be obtained in real-time by reading the encoder values ​​of each joint of the robotic arm and combining them with a forward kinematics algorithm. Alternatively, it can be obtained by measuring optical markers mounted on the gripper surface and by an external motion capture system. It can also be obtained by reading the rigid body pose from the physical interface of the simulation engine in a robot simulation device, serving as the single gripper pose information. Then, the single gripper pose information is mapped to first-order pose data, serving as the spatial reference for subsequent calculations of the target's relative pose.

[0040] The opening and closing degree of a single gripper characterizes the instantaneous state of physical interaction between the actuator and the object, and its numerical form can be flexibly defined according to the sensor type. For example, for an electric gripper, it can be a dimensionless value normalized to the interval [0, 1], where 0 represents fully closed and 1 represents fully open; it can also be the original pulse count of the motor encoder or the joint angle value; for pneumatic or hydraulically driven grippers, it can be the reading of a pneumatic / hydraulic sensor or the opening percentage of a proportional valve.

[0041] In the semantic space of imitation learning, position and gesture, as well as grasping force, belong to two separate feature dimensions. For example, gesture indicates where the hand is, while the degree of opening indicates what the hand is doing. Therefore, the degree of opening is defined as the first historical action data rather than being incorporated into the gesture information. This decouples the two feature data, gesture and grasping force, allowing the model to more accurately identify grasping or releasing actions triggered at specific spatial locations, thus avoiding feature confusion caused by the coupling of the two feature data.

[0042] It is understood that this embodiment only uses a real single gripper in the physical world as an example for illustration. The pose and opening / closing state of the single gripper can also be the pose and opening / closing state data of a virtual single gripper in a digital twin or pure simulation training scenario.

[0043] Through the above methods, this embodiment defines a specific data structure for single-arm or single-hand operation scenarios, constructs an input feature space for single-arm imitation learning tasks, clarifies the mapping relationship between pose and motion data under single gripper conditions, and improves the model's recognition accuracy for triggering grasping or releasing actions at specific spatial positions.

[0044] As one implementation, the method further includes: When the execution end is a dual gripper, the left gripper pose information and the right gripper pose information of the execution end at the first moment are collected, and the left gripper pose information and the right gripper pose information are combined to obtain the first pose data; The opening and closing degrees of the left and right grippers at the first moment are collected from the execution end, and the opening and closing degrees of the left and right grippers are combined to obtain the first historical action data. In some embodiments, a dual gripper refers to a configuration with two independent or cooperative end effectors, including but not limited to a humanoid robot's dual-arm system, an industrial dual-arm collaborative robot, or two data acquisition devices held by an operator in each hand. To simultaneously characterize the spatial state and interaction intent of the two ends in a dual-gripper scenario, a specific combination logic is determined based on the spatial semantics and temporal relationships necessary for hand-to-hand collaboration, integrating the data from both sides into a unified model input feature.

[0045] Specifically, the combination of the left gripper pose information and the right gripper pose information can be achieved through vector concatenation, where the six-DOF pose vectors of the left and right grippers are concatenated end-to-end in a preset order to obtain a joint pose vector as the first pose data; it can also be achieved through a combination of structured dictionaries or key-value pairs, such as constructing a composite object containing left_pose and right_pose fields; or it can be achieved through weighted fusion, where the poses of the left and right grippers are weighted and fused based on preset weights or attention mechanisms to obtain a representative center pose as the first pose data; thus, the bilateral states corresponding to the two grippers are mapped to a unified input tensor that can be recognized by the imitation learning model.

[0046] The combination of the opening and closing degrees of the left and right grippers can be achieved by concatenating their scalar values ​​into a two-dimensional vector, or by encapsulating them into a composite feature containing the grasping states of both sides, thus obtaining the first historical motion data. The combined motion data, while representing the grasping intent of each end effector, also includes the mechanical constraints of the coordinated operation of both hands, such as the need for both hands to close synchronously when carrying large objects, or the differentiated opening and closing patterns of one hand fixing and the other rotating in assembly tasks.

[0047] By employing the above method, the data in dual-gripper scenarios is input into the execution-end state model as a whole. This allows the model to simultaneously learn the spatial layout and collaborative relationship of both hands relative to the head's perspective, rather than predicting the movements of each hand in isolation. For example, when the model receives the combined first pose data, it can perceive the positional shift of the left hand relative to the right hand, thereby accurately inferring whether the current action is a two-handed grasping, passing, or independent operation, thus improving the success rate of imitating complex dual-arm tasks.

[0048] As one implementation method, the model training method further includes: Based on the second pose and the pose information of the left gripper, calculate the first pose transformation matrix of the view observation component relative to the left gripper; Based on the second pose and the pose information of the right gripper, calculate the second pose transformation matrix of the visual observation component relative to the right gripper; The target relative pose data is obtained based on the first pose transformation matrix and the second pose transformation matrix. In some embodiments, in a dual-arm collaborative operation scenario, the left and right grippers are typically in different positions and postures in space, and the operator's head viewpoint relative to the left and right hands constitutes different spatial geometric relationships. For example, one hand may be holding a workpiece while the other is tightening screws, or both hands may be used to collaboratively move large, irregularly shaped objects. If the combination method described in the above embodiments is used, that is, the poses of the left and right grippers are simply spliced ​​together and a unified overall relative pose is calculated, although bilateral information is preserved at the data structure level, the spatial semantics of each hand relative to the observation viewpoint are lost at the geometric calculation level. Therefore, this embodiment calculates two independent pose transformation matrices based on the differentiated viewpoints and arm relationships, providing the model with a rich description of the spatial relationship between the head and hands.

[0049] Specifically, the second pose of the visual observation component and the pose information of the left gripper are obtained. The second pose represents the spatial pose of the visual observation component in its own coordinate system or a certain global reference coordinate system, while the left gripper pose information represents the spatial pose of the left gripper in the execution end coordinate system or the same global reference coordinate system. Based on the second pose and the pose information of the left gripper, the spatial pose relationship between the visual observation component and the left gripper is obtained through coordinate transformation operations, i.e., the first pose transformation matrix. This first pose transformation matrix is ​​a 4×4 homogeneous transformation matrix, containing a 3×3 rotation submatrix and a 3×1 translation vector, used to describe the six-degree-of-freedom pose of the left gripper in the coordinate system of the visual observation component, or to describe the spatial relationship between the visual observation component and the left gripper. For example, when the second pose is represented by the homogeneous transformation matrix T_cam in the world coordinate system, and the left gripper pose information is represented by the homogeneous transformation matrix T_left in the world coordinate system, the first pose transformation matrix can be obtained by multiplying the inverse matrix of T_cam with T_left, that is, T_left_cam = inv(T_cam)×T_left. This matrix is ​​the pose expression of the left gripper in the coordinate system of the visual observation component.

[0050] Based on the second pose of the visual observation component and the pose information of the right gripper, the spatial pose relationship between the visual observation component and the right gripper is obtained through coordinate transformation operations, namely, T_right_cam = inv(T_cam)×T_right, where T_right is the pose transformation matrix of the right gripper in the world coordinate system. This matrix is ​​used to describe the spatial state of the right gripper from the head's perspective. Through the above matrix operations, two pose transformation matrices corresponding to the left and right grippers are obtained respectively. Each of the two matrices includes complete six degrees of freedom information of one arm relative to the head's perspective, avoiding spatial semantic confusion that may be introduced due to data combination.

[0051] After obtaining the first and second pose transformation matrices, the target relative pose data is obtained based on them. The target relative pose data can be combined by concatenating the first and second pose transformation matrices in a preset order into a joint matrix or joint vector, which serves as the final target relative pose data. For example, two 4×4 homogeneous transformation matrices can be flattened and concatenated end-to-end to form a joint vector; or the translation components and rotation quaternions from the two matrices can be extracted separately and concatenated into a feature vector. This ensures that the poses of both grippers are expressed in the common reference frame corresponding to the visual observation component, eliminating numerical deviations caused by coordinate system inconsistencies. The target relative pose data can also be combined by deriving the relative spatial relationship between the two arms based on the first and second pose transformation matrices. For example, the pose transformation of the left gripper relative to the right gripper, T_left_right = inv(T_left_cam)×T_right_cam, can be calculated, and this relative pose between the arms can be used as a supplementary dimension of the target relative pose data, enabling the model to impose geometric constraints on the cooperation between the two hands.

[0052] This embodiment preserves the spatial semantic integrity of each arm relative to the observation viewpoint by separately calculating and combining two pose transformation matrices. This allows the model to not only identify the precise positions and postures of the left and right hands from the head's perspective (rather than just estimating the overall, ambiguous position of both arms), but also to identify asymmetric operations that may occur in dual-arm tasks. For example, when one hand is at the edge of the workspace while the other is operating in the central area, the difference in their viewpoints relative to the head is significant. Calculating two pose transformation matrices separately can accurately capture this difference.

[0053] By employing the above method, this embodiment calculates the independent transformation matrix of the head relative to each gripper in a dual-arm scenario, preserving the independent spatial semantics of each hand and avoiding confusion in the spatial relationship between the left and right hands caused by simple splicing. This facilitates the model's refined learning of the relative positional constraints in dual-arm operations. Consequently, it provides richer feature input dimensions for model training, enabling the model to simultaneously receive the joint pose of both arms at the global level and the independent relative pose of each arm at the local level, thereby learning the strategy rules of dual-arm operations at different levels.

[0054] It is understood that the transformation matrix calculation strategy of this embodiment is also applicable to multi-finger dexterous hands or other multi-end effector configurations containing more end effectors. For example, for an actuator containing N independent controllable ends, N corresponding pose transformation matrices can be calculated in parallel and combined into target relative pose data.

[0055] As one implementation, step S102 includes: The first pose data, the first historical action data, the historical visual observation data, and the target relative pose data are used as input features; Based on the second historical action data, the output label is used, and the imitation learning model is trained based on the input features and the output label. In some embodiments, historical visual observation data typically includes RGB image sequences, depth maps, or point clouds, which are processed by a visual encoder such as a convolutional neural network or a Vision Transformer to extract high-level semantic features. The first pose data and the first historical action data are used to characterize the spatial position and interaction intent of the execution end, respectively, and can be embedded through a multilayer perceptron or a linear layer. The target relative pose data serves as a bridge connecting visual perception and ontological control and is fused into the feature stream. By training the model with the target relative pose data, the model can acquire accurate spatial relationships between the head and hands in each forward propagation, avoiding illusions or error accumulation caused by the model relying on unstable visual features to infer scale.

[0056] It is understood that, in addition to the feature splicing method shown in the above embodiments, input features can also be spliced ​​using methods such as cross-attention mechanism, gated fusion unit or tensor product.

[0057] The second historical motion data represents the ideal motion that the actuator should perform in the next moment or within a future time window. It can take the form of end effector pose increments, joint angle sequences, gripper opening and closing commands, or torque control signals, and is used as the output label. During training, the model predicts a motion sequence based on the current input features, then calculates the difference between the predicted motion and the second historical motion data (which serves as the ground truth), and updates the model parameters using the backpropagation algorithm.

[0058] Specifically, for translational components or scalar control signals (such as gripper opening degree), L2 loss (mean squared error) or Smooth L1 loss can be used to measure the deviation between the predicted value and the label; for rotational components, geodesic distance loss or quaternion cosine similarity loss is used to ensure that the calculation of rotational error conforms to the geometric nature of three-dimensional space. Different weighting coefficients can also be assigned to each sub-item of loss according to task requirements; for example, increasing the weight of rotational loss in assembly tasks, while increasing the weight of translational loss in large-scale handling tasks.

[0059] By using the above method, the target's relative pose, other ontological states, and visual observations are used as input features, enabling the model to predict future actions based on the current spatial relationship between the head and hands, thus improving the model's accuracy in perceiving spatial scale from a first-person perspective. Please see Figure 3 , Figure 3This is a schematic flowchart illustrating another model training method provided in an embodiment of this application. Figure 3 As shown, the method further includes: Step S103: Based on the historical relative pose data at each time point in the historical sample data, calculate the historical pose change between the historical relative pose data at each adjacent time point. The historical relative pose data is the relative pose data between the visual observation component and the execution end. Step S104: Based on the comparison result of the historical pose change amount and the preset threshold, determine the relevant time when the execution end moves itself in the historical sample data; Step S105: Select at least one of the relevant times as the first time. In some embodiments, effective interaction segments are determined from massive time-series data by dynamically filtering pose changes, thereby improving training efficiency and model performance.

[0060] In practical implementation, the historical pose change includes two dimensions: the change in spatial position and the deflection of the attitude angle. For the translation component, Euclidean distance can be used to calculate the change in the relative position between the visual observation component and the execution end between two adjacent frames; for the rotation component, geodesic distance or quaternion angle can be used to characterize the degree of change in relative attitude. Judgment conditions can be set separately for displacement and rotation, or different weights can be assigned according to the task characteristics before calculating the comprehensive change. For example, in the acquisition of assembly tasks, the displacement threshold can be set to 1 cm and the rotation threshold to 5 degrees; while in large-scale handling tasks, the displacement threshold can be relaxed to 3 cm and the rotation threshold adjusted to 10 degrees.

[0061] Understandably, the specific value of the aforementioned threshold can be set based on the noise level of the positioning sensor, the accuracy requirements of the operational task, and the frequency of data acquisition. For example, when the sensor has high-frequency jitter noise, the threshold can be increased to avoid misjudging noise as valid motion; or preprocessing methods such as sliding window averaging and low-pass filtering can be used to smooth the data before comparison to enhance the robustness of the screening results. The relevant moments determined based on the above comparison results are the time points at which the execution end undergoes a substantial spatial transformation relative to the visual observation component. Using these moments as the first moment for model training ensures that the model updates parameters only on data points with clear dynamic intent. This not only reduces the number of invalid samples but also allows the model to learn more discriminative spatiotemporal features during critical windows of intense relative movement between the head and hand.

[0062] Even if the execution device moves at high speed in the world coordinate system, if it remains relatively stationary with respect to the head (such as when a handheld camera follows an object), its contribution to operational skills is far lower than when the head and hand are relatively misaligned. Therefore, compared with manual annotation or screening methods based on simple speed thresholds, the screening based on the relative pose change of the head and hand provided in this embodiment can more accurately reflect the interactive semantics in the first-person perspective.

[0063] It is understood that this embodiment only uses the selected relevant time as the first time as an example for illustration. In other embodiments, the relevant time can also form a candidate sampling pool, and then perform secondary sampling from it according to a uniform distribution or a specific probability distribution, so as to further balance data diversity and training stability.

[0064] In this embodiment, key training samples are screened by monitoring the dynamic changes in the relative poses of the head and hands, filtering out static or invalid motion segments, thereby increasing the effective density of training data and improving model training efficiency. As one implementation, the method further includes: Based on the positioning module on the visual observation component, the second pose of the visual observation component is acquired, and the positioning module is integrated on the visual observation component. In some embodiments, the integration of the positioning module into the visual observation component means that the two form an inseparable whole in terms of physical structure or a fixed relative position.

[0065] Specifically, the positioning module can be embedded inside the housing of the visual observation component, sharing the same main control circuit board as the image sensor; alternatively, it can be rigidly connected to the outer surface of the visual observation component via a mechanical interface, such as by screws, clips, or adhesives, to the top or side of the camera housing. By integrating the positioning module onto the visual observation component, a definite and stable spatial transformation relationship is established between the measurement reference point of the positioning module and the optical center of the visual observation component.

[0066] It is understood that, in addition to the head-mounted device form shown in the above embodiments, the visual observation component can also be a camera module mounted on the robot's head, mobile chassis, or a third-party bracket.

[0067] In specific embodiments, the positioning module may employ an inertial measurement unit (IMU) to calculate pose in real time by measuring triaxial acceleration and angular velocity and combining them with an integral algorithm; it may also include an ultra-wideband (UWB) tag or a Bluetooth AOA positioning unit to obtain absolute position information by interacting with preset base stations in the environment; it may also be a passive optical marker or an active infrared light-emitting diode array, used in conjunction with an external motion capture camera for tracking; or it may be a dedicated computing module integrating a visual odometry (VIO) algorithm, using its own camera to assist inertial navigation for drift correction; or it may employ a multi-sensor fusion scheme, such as combining an IMU with UWB or VIO.

[0068] In this embodiment, by rigidly connecting the positioning module and the camera, the relative pose (i.e., hand-eye calibration parameters) between the two remains constant, which not only ensures the accuracy of spatiotemporal synchronization from the source of data acquisition, but also simplifies the on-site deployment process.

[0069] In one embodiment, the visual observation component is disposed at the front of a head-mounted bracket, which is used to fix the visual observation component to the operator's head. The optical axis of the visual observation component is aligned with the operator's line of sight, and the visual observation component is used to collect visual observation data from a first-person perspective. In some embodiments, the visual observation component may be a head-mounted device, such as... Figure 4 As shown, 1 is the visual observation component, and 2 is the head-mounted support. The positioning module can be integrated into the visual observation component. The head-mounted support fixes the visual observation component to the operator's head area and keeps it in a relatively fixed posture, ensuring that the visual observation component will not slip or fall off when the operator makes large limb movements or moves quickly.

[0070] Understandably, head-mounted braces can be like... Figure 4 The headband-style support shown can also be a helmet liner, an AR / VR glasses frame, a hat brim clip, or a bone conduction headphone-style ear hook structure.

[0071] Specifically, the optical axis direction of the visual observation component aligning with the operator's line of sight means that the optical axis direction of the visual observation component should be parallel to or have only a slight angle (e.g., less than 10 degrees) with the normal vector of the line connecting the operator's eyes, and its field of view (FOV) should cover the operator's primary area of ​​focus when performing the task. Therefore, the visual observation component is positioned in the central front region of the head-mounted device, corresponding to the horizontal line at eye level.

[0072] In a specific embodiment, to accommodate differences in facial contours and wearing habits among different operators, a multi-dimensional fine-tuning mechanism can be provided between the visual observation component and the head-mounted support, such as a pitch adjustment knob, a horizontal translation groove, or a ball joint. This fine-tuning mechanism allows the operator to quickly calibrate the camera's viewing angle after each use, ensuring it precisely aligns with the operator's natural gaze direction. This ensures that the collected historical viewing angle observation data accurately reflects the visual input characteristics used during task execution.

[0073] It is understandable that, in addition to being a head-mounted device, the visual observation component can also be fixed to the operator's chest, shoulders, or waist. The visual observation component can also be mounted on a ring-shaped fixed bracket, which includes a tripod, a ceiling bracket, or a wall bracket, and the bracket can be determined according to the operator's viewing angle. The visual observation component can be a monocular RGB camera, a binocular stereo camera, an RGB-D depth camera, or a multispectral sensor module. Specifically, multiple auxiliary cameras can be symmetrically arranged on the head-mounted bracket to expand the peripheral field of view.

[0074] By fixing the visual observation component to the operator's head using the above method and a head-mounted bracket, the collected visual data is highly aligned with the operator's natural line of sight, thereby maximizing the reproduction of the user's real perceptual experience when performing the task and providing first-person perspective observation data for the imitation learning model.

[0075] Please see Figure 5 , Figure 5 This is a schematic block diagram of a robot device provided in an embodiment of this application.

[0076] See Figure 5 The robot includes: The visual observation component is used to collect visual observation data of the execution end when it performs actions. The first positioning module is used to locate the pose data of the execution end; The second positioning module is disposed on the visual observation component and is used to locate the pose data of the visual observation component; The processor, memory, and network interface are connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0077] The non-volatile storage medium can store the operating system and computer program. The computer program includes program instructions, which, when executed, cause the processor to execute the computer program to implement the steps of the model training method described above.

[0078] The processor provides computing and control capabilities to support the operation of the entire robot.

[0079] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, the processor can execute the computer program to implement the steps of the above-mentioned model training method.

[0080] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the robot to which the present application is applied. A specific robot may include more or fewer parts than shown in the figure, or combine certain parts, or have different part arrangements.

[0081] It should be understood that the processor can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is used to run a computer program stored in memory to implement the steps of the model training method provided in the embodiments of this application.

[0082] The computer-readable storage medium can be the internal storage unit of the robot described in the foregoing embodiments, such as the robot's hard drive or memory. Alternatively, the computer-readable storage medium can be an external storage device for the robot, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card.

[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A model training method, characterized in that, The model training method includes: Based on the first pose of the execution end and the second pose of the visual observation component at the first moment, the target relative pose is calculated, and the target relative pose represents the spectral spatial relationship between the visual observation component and the execution end at the first moment. Based on the first pose data, the target relative pose data, the historical visual observation data, the first historical action data of the execution end at the first moment, and the second historical action data at the next moment, the imitation learning model is trained based on the localization of the observation component; wherein, the historical visual observation data is collected by the visual observation component.

2. The method as described in claim 1, characterized in that, The model training method also includes: When the first pose data and the second pose data belong to different coordinate systems, the first pose data and the second pose data are transformed to the same target coordinate system based on preset calibration parameters; Based on the first pose data and the second pose data in the target coordinate system, the pose transformation of the visual observation component relative to the execution end is calculated, which is used as the target relative pose data.

3. The method as described in claim 2, characterized in that, The step of converting the first pose data and the second pose data to the same target coordinate system based on preset calibration parameters includes: The coordinate system of the visual observation component to which the second pose data belongs is used as the target coordinate system; Based on the transformation matrix corresponding to the first pose data and the second pose data, the first pose data is transformed from the execution end coordinate system to the visual observation component coordinate system.

4. The method as described in claim 1, characterized in that, The method further includes: When the execution end is a single gripper, the single gripper pose information of the execution end at the first moment is collected as the first pose data; The opening and closing degree of the single gripper at the first moment of the execution end is collected as the first historical action data.

5. The model training method as described in claim 1, characterized in that, The method further includes: When the execution end is a dual gripper, the left gripper pose information and the right gripper pose information of the execution end at the first moment are collected, and the left gripper pose information and the right gripper pose information are combined to obtain the first pose data; The opening and closing degrees of the left and right grippers at the first moment are collected from the execution end, and the opening and closing degrees of the left and right grippers are combined to obtain the first historical action data.

6. The model training method as described in claim 5, characterized in that, The model training method also includes: Based on the second pose and the pose information of the left gripper, calculate the first pose transformation matrix of the view observation component relative to the left gripper; Based on the second pose and the pose information of the right gripper, calculate the second pose transformation matrix of the visual observation component relative to the right gripper; The target relative pose data is obtained based on the first pose transformation matrix and the second pose transformation matrix.

7. The model training method as described in claim 1, characterized in that, The training of the imitation learning model based on observation part localization, using first historical action data, first pose data, second historical action data, second pose data, target relative pose data, and historical visual observation data, includes: The first pose data, the first historical action data, the historical visual observation data, and the target relative pose data are used as input features; Based on the second historical action data, the output label is used, and the imitation learning model is trained based on the input features and the output label.

8. The model training method as described in claim 1, characterized in that, The method further includes: Based on the historical relative pose data at each time point in the historical sample data, the historical pose change between the historical relative pose data at each adjacent time point is calculated. The historical relative pose data is the relative pose data between the visual observation component and the execution end. Based on the comparison results between the historical pose change and the preset threshold, the relevant moments when the execution end moves are determined in the historical sample data. At least one of the relevant times is taken as the first time.

9. The model training method as described in claims 1-8, characterized in that, The method further includes: Based on the positioning module on the visual observation component, the second pose data of the visual observation component is collected, and the positioning module is integrated on the visual observation component.

10. The model training method as described in claim 9, characterized in that, The visual observation component is disposed at the front of the head-mounted bracket, which is used to fix the visual observation component to the operator's head. The optical axis of the visual observation component is aligned with the operator's line of sight. The visual observation component is used to collect visual observation data from a first-person perspective.

11. A robot, characterized in that, include: The visual observation component is used to collect visual observation data of the execution end when it performs actions. The first positioning module is used to locate the pose data of the execution end; The second positioning module is disposed on the visual observation component and is used to locate the pose data of the visual observation component; A processor, a memory, and an action generation program based on observation component localization stored in the memory and executable by the processor, wherein the action generation program based on observation component localization, when executed by the processor, implements the steps of the model training method as described in any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an action generation program based on the localization of the observation component, wherein when the action generation program based on the localization of the observation component is executed by a processor, it implements the steps of the model training method as described in any one of claims 1 to 10.