Model-free robot control method and system based on object pose trajectory
Patent Information
- Application Number
- CN202611096551.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-23
AI Technical Summary
[0003]现有技术通常将演示生成、物体感知与机器人控制作为相互独立的功能模块进行组合,缺乏以统一中间表示贯穿全流程的一体化技术方案,因而在实际应用中仍存在对人工示教依赖较强、对未知物体适应能力不足、跨形态迁移困难以及难以实现闭环控制的问题
本发明首先,基于所述初始视觉观测信息与所述任务语义信息生成候选任务演示视频;在候选任务演示视频中筛选得到描述目标物体期望运动过程的目标演示视频;然后,基于所述初始视觉观测信息与所述目标演示视频,执行免模型位姿估计,以获取目标物体在连续时间尺度下的六自由度位姿轨迹;以所述六自由度位姿轨迹作为统一中间表示,在仿真环境中对机器人策略进行训练,确定执行策略;最后,基于所述执行策略控制机器人执行任务,并获取物体当前位姿;基于物体当前位姿与所述六自由度位姿轨迹之间的偏差对机器人动作进行闭环调整。通过构建视觉序列、三维重建和时序位姿优化的统一轨迹解析机制,将生成视频中的隐式动作信息转化为具有物理一致性的物体六自由度位姿轨迹,实现了跨模态数据的统一表达,在此基础上,以物体位姿轨迹为唯一共享约束变量,将其同时作为感知输出、仿真学习目标及真实执行反馈基准,构建跨阶段一致的优化目标,并通过闭环反馈机制实现轨迹偏差的在线修正,从而实现从任务生成到机器人执行的统一融合,避免了简单叠加式集成,形成端到端一致性的技术方案,从架构层面提升了系统整体一致性与泛化能力,同时无需依赖人工示教、物体三维模型及人体姿态检测,显著降低了数据获取成本与系统部署复杂度。
Smart Images

Figure CN122584377B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of robot vision perception and operation control technology, and in particular relates to a model-free robot control method and system based on object pose trajectory. Background Technology
[0002] Imitation learning for robotic arm operations is a core research problem in the field of robot perception and control, and it is widely used in industrial automation, flexible manipulation, and service robots. Compared to traditional teaching programming, imitation learning can significantly reduce deployment costs and improve the robot's generalization ability in unstructured environments by replicating the motion patterns of objects during operations. In the complete chain of imitation learning, how to connect task understanding, environmental perception, and policy control in a unified way is crucial to the successful implementation of the system.
[0003] Existing technologies typically combine demonstration generation, object perception, and robot control as independent functional modules, lacking an integrated technical solution with a unified intermediate representation throughout the entire process. As a result, in practical applications, there are still problems such as strong reliance on human teaching, insufficient adaptability to unknown objects, difficulty in cross-morphological transfer, and difficulty in achieving closed-loop control. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes a model-free robot control method and system based on object pose trajectory. By constructing a unified trajectory parsing mechanism integrating visual sequence, 3D reconstruction, and temporal pose optimization, the implicit motion information in the generated video is transformed into a physically consistent six-DOF object pose trajectory, achieving a unified expression of cross-modal data. Based on this, the object pose trajectory is used as the sole shared constraint variable, simultaneously serving as the perception output, simulation learning target, and real execution feedback benchmark. This constructs a consistent optimization target across stages, and online correction of trajectory deviations is achieved through a closed-loop feedback mechanism. This unifies the process from task generation to robot execution, avoiding simple additive integration and forming an end-to-end consistent technical solution. From an architectural perspective, this improves the overall consistency and generalization capability of the system. Furthermore, it eliminates the need for manual teaching, object 3D models, and human pose detection, significantly reducing data acquisition costs and system deployment complexity.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solution: In a first aspect, the present invention provides a model-free robot control method based on object pose trajectory, comprising: Acquire initial visual observation information and task semantic information; Candidate task demonstration videos are generated based on the initial visual observation information and the task semantic information; target demonstration videos describing the expected motion process of the target object are selected from the candidate task demonstration videos. Based on the initial visual observation information and the target demonstration video, model-free pose estimation is performed to obtain the six-degree-of-freedom pose trajectory of the target object in a continuous time scale; the six-degree-of-freedom pose trajectory is used as a unified intermediate representation to train the robot strategy in a simulation environment and determine the execution strategy. The robot is controlled to perform tasks based on the execution strategy and the current pose of the object is obtained; the robot's actions are adjusted in a closed loop based on the deviation between the current pose of the object and the six-degree-of-freedom pose trajectory.
[0006] Furthermore, the model-free pose estimation includes: constructing point-level motion information of the object by predicting the two-dimensional point correspondence between the starting frame and the target frame, and recovering the three-dimensional point coordinates by combining the depth information; and estimating the six-degree-of-freedom pose trajectory of the object based on the three-dimensional point coordinates.
[0007] Furthermore, the recovery of the three-dimensional point coordinates includes: using the initial observation image as the anchor frame, performing frame-by-frame matching on each frame of the generated video, constructing cross-frame correspondence by combining appearance features and geometric constraints; after obtaining a stable two-dimensional pixel correspondence, combining the depth information and camera intrinsic parameters in the initial observation, back-projecting the two-dimensional pixel points to three-dimensional space, recovering the corresponding three-dimensional point coordinates, thereby constructing a three-dimensional point set representation of the target object.
[0008] Furthermore, based on the aforementioned 3D point set representation, the initial pose estimate of the target object in the current frame is first obtained by performing global geometric matching on the 3D point set; then, the pose parameters are iteratively optimized by minimizing the reprojection error between the observation point in the current frame and the reference point set.
[0009] Furthermore, the initial visual observation information is an RGB-D image, and the task semantic information is a natural language instruction.
[0010] Furthermore, the generation of the candidate task demonstration video includes: based on a temporal generation network, using initial visual observation information as scene content constraints and task semantic information as behavioral condition input, mapping language instructions to visual action priors through cross-modal conditional encoding, guiding the model to generate dynamic video sequences that conform to task semantics; in the process of generating dynamic video sequences, multiple sets of candidate video sequences are obtained by fixing the conditions of the initial frame and applying consistency constraints in the temporal dimension.
[0011] Furthermore, the determination of the target demonstration video includes: for each candidate video in the candidate video sequence, using semantic consistency as a prior constraint to filter out sequences that do not meet the task objective, and then performing joint optimization and filtering based on object identity consistency and physical rationality.
[0012] Furthermore, the closed-loop adjustment includes: during execution, determining the position deviation and attitude deviation based on the difference between the real-time estimated object pose and the target trajectory, and maintaining the current control strategy when the deviation is within a preset range.
[0013] Furthermore, when the deviation exceeds the first threshold, a local online adjustment mechanism is triggered, introducing a feedback correction term based on the current strategy output to incrementally compensate the robot's end effector motion, so that the object gradually returns to the target trajectory; when the deviation exceeds the second threshold, a backtracking and replanning mechanism is triggered, backtracking the state to the most recent historical node that met the stability conditions, and regenerating the subsequent execution action sequence based on the historical node.
[0014] Secondly, the present invention also provides a model-free robot control system based on object pose trajectory, comprising: The data acquisition module is configured to acquire initial visual observation information and task semantic information; The demonstration video determination module is configured to: generate candidate task demonstration videos based on the initial visual observation information and the task semantic information; and select target demonstration videos that describe the expected motion process of the target object from the candidate task demonstration videos. The execution strategy determination module is configured to: perform model-free pose estimation based on the initial visual observation information and the target demonstration video to obtain the six-degree-of-freedom pose trajectory of the target object in a continuous time scale; use the six-degree-of-freedom pose trajectory as a unified intermediate representation to train the robot strategy in a simulation environment and determine the execution strategy; The control module is configured to: control the robot to perform tasks based on the execution strategy and obtain the current pose of the object; and perform closed-loop adjustment of the robot's actions based on the deviation between the current pose of the object and the six-degree-of-freedom pose trajectory.
[0015] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the model-free robot control method based on object pose trajectory described in the first aspect.
[0016] Fourthly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the program to implement the steps of the model-free robot control method based on object pose trajectory described in the first aspect.
[0017] Fifthly, the present invention also provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the steps of the model-free robot control method based on object pose trajectory described in the first aspect.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention first generates candidate task demonstration videos based on the initial visual observation information and the task semantic information; then, it filters the candidate task demonstration videos to obtain a target demonstration video describing the expected motion process of the target object; next, based on the initial visual observation information and the target demonstration video, it performs model-free pose estimation to obtain the six-degree-of-freedom pose trajectory of the target object over a continuous time scale; using the six-degree-of-freedom pose trajectory as a unified intermediate representation, it trains the robot strategy in a simulation environment to determine the execution strategy; finally, it controls the robot to perform the task based on the execution strategy and obtains the current pose of the object; and it performs closed-loop adjustment of the robot's actions based on the deviation between the current pose of the object and the six-degree-of-freedom pose trajectory. By constructing a unified trajectory parsing mechanism for visual sequence, 3D reconstruction, and temporal pose optimization, implicit motion information in generated videos is transformed into physically consistent six-DOF object pose trajectories, achieving a unified expression of cross-modal data. Based on this, the object pose trajectory is used as the only shared constraint variable, simultaneously serving as the perception output, simulation learning target, and real execution feedback benchmark, constructing a consistent optimization target across stages. A closed-loop feedback mechanism is used to achieve online correction of trajectory deviations, thereby achieving unified integration from task generation to robot execution. This avoids simple superposition integration, forming an end-to-end consistent technical solution, improving the overall consistency and generalization capability of the system at the architectural level. At the same time, it eliminates the need for manual teaching, object 3D models, and human pose detection, significantly reducing data acquisition costs and system deployment complexity. Attached Figure Description
[0019] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0020] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the model-free six-DOF pose estimation network and visual motion strategy network structure based on object pose trajectory driven in Embodiment 1 of the present invention. Figure 3 This is a schematic diagram of the simulation environment for Embodiment 1 of the present invention. Detailed Implementation
[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0023] Example 1: Imitation learning for robotic arm operations is a core research problem in the field of robot perception and control, and it is widely used in industrial automation, flexible manipulation, and service robots. Compared with traditional teaching programming, imitation learning can significantly reduce deployment costs and improve the robot's generalization ability in unstructured environments by replicating the motion patterns of objects during operations.
[0024] In the complete chain of imitation learning, how to connect task understanding, environmental perception, and policy control in a unified way is the key to determining whether the system can be implemented. Existing technologies generally suffer from problems such as fragmented links, independent modules, and a lack of unified intermediate representation: task demonstrations rely on manual data collection, perception and control are independent of each other, and trajectory information cannot form a consistent constraint between generation, perception, and control, making it difficult for the system to form an organic whole.
[0025] From a perception perspective, traditional solutions heavily rely on object CAD models, multi-view reconstruction, or manual pre-calibration, making them unable to handle unknown objects without prior information. Although model-free pose estimation methods exist, they are mostly limited to single-frame inference and cannot provide continuous, stable, and controllable object trajectories. Existing trajectory tracking also suffers from insufficient initialization accuracy, temporal drift, and disconnection from the control objective.
[0026] From a control perspective, existing imitation learning methods mostly use human joint trajectories or end-effector trajectories as supervision, which leads to prominent contradictions with the differences between human and machine structures and the mismatch of workspaces. At the same time, there is a lack of a unified interface between trajectory data and simulation transfer and real execution, resulting in a large gap between simulation and the real domain, poor policy generalization, and insufficient closed-loop robustness.
[0027] In summary, existing technologies typically combine demonstration generation, object perception, and robot control as independent functional modules, lacking an integrated technical solution with a unified intermediate representation throughout the entire process. As a result, in practical applications, there are still problems such as strong reliance on human teaching, insufficient adaptability to unknown objects, difficulty in cross-morphological transfer, and difficulty in achieving closed-loop control.
[0028] To address at least one of the aforementioned problems, this embodiment provides a model-free robot control method based on object pose trajectory. The method uses object pose trajectory as a unified intermediate representation to connect the generation, perception, and control processes, enabling closed-loop execution and cross-scenario generalization of robot manipulation tasks without the need for manual demonstration or prior models.
[0029] This embodiment uses the object's pose trajectory as a unified intermediate representation connecting task understanding, visual perception, and robot control. It integrates generative visual demonstration, model-free pose estimation, cross-morphological simulation learning, and closed-loop execution in a real-world environment, rather than simply combining them as independent modules. By introducing unified pose trajectory constraints, it achieves collaborative optimization of each processing stage, enabling the system to complete closed-loop execution of complex tasks without manual demonstration, prior object models, or manual annotation, and possesses good cross-scene adaptability and robustness. Figure 1 As shown, the method in this embodiment includes: S1. Obtain initial observations and generate the task demonstration sequence corresponding to the target trajectory: Acquire initial visual observation information and task semantic information in a real environment, wherein the initial visual observation information is a single-frame RGB-D image and the task semantic information is a natural language instruction.
[0030] The RGB-D image is preprocessed, including target object region extraction, depth alignment, and camera parameter calibration, to obtain a unified initial observation representation.
[0031] Based on the initial observation information and task semantic information, a conditional visual generation model is used to generate candidate video sequences corresponding to the target task. Specifically, the conditional visual generation model is a temporal generation network based on a diffusion mechanism. It uses initial visual observation information as scene content constraints and task semantic information as behavioral condition input. Through cross-modal conditional encoding, it maps language instructions to visual action priors, guiding the model to generate dynamic video sequences that conform to the task semantics. During the generation process, by fixing the initial frames and applying consistency constraints in the temporal dimension, the generated video gradually evolves into an object motion process that conforms to the task objective while maintaining consistency between the scene structure and the appearance of the target object, thus obtaining multiple sets of candidate video sequences. Furthermore, diverse candidate results are generated through multiple random samplings to improve task coverage and generation robustness. These candidate video sequences implicitly represent the desired motion trajectory of the target object.
[0032] Multi-constraint joint screening is performed on the candidate video sequences to obtain target demonstration sequences for trajectory parsing. Specifically, for each candidate video, a comprehensive evaluation is conducted from three dimensions: semantic consistency, object identity consistency, and physical plausibility. Regarding semantic consistency, the task semantic information is converted into text feature vectors, keyframe sequences in the candidate videos are encoded into visual feature vectors, and input into a visual-language alignment model for cross-modal matching. A semantic consistency score is obtained by calculating the similarity between action execution and instruction description. ; in, The text feature vector representing the semantic information of the task. This represents the visual feature vector corresponding to the candidate video. The semantic consistency score is represented by the number of points. A higher score indicates that the action process expressed in the video is more consistent with the task semantics. Regarding object identity consistency, feature matching and instance-level association are performed on target objects in the initial observation image and video frames to constrain the continuity of the target object's appearance features and spatial position in the temporal sequence, thus preventing object replacement or drift. Regarding physical plausibility, the velocity, acceleration, and contact relationship changes of objects in the video are estimated based on their motion trajectories, and abnormal motion is detected through preset physical constraint rules to eliminate candidate sequences that do not conform to real physical laws.
[0033] Furthermore, to reflect the synergistic constraints among semantic consistency, object identity consistency, and physical rationality, this embodiment employs a phased joint screening mechanism, rather than independently thresholding each indicator. First, the semantic consistency score is... As a priori constraint index. When Below the preset semantic threshold If a candidate video does not match the task objective, it is directly determined and eliminated; only candidate videos that meet the semantic constraints are retained for the next stage of screening. For candidate videos that pass the semantic screening, identity consistency scores are further considered. Physical rationality score Construct a joint scoring function: ; in, β and γ These are weights for identity consistency and physical plausibility, respectively. Since semantic consistency scores participate in the overall evaluation multiplicatively, when a candidate video deviates from the task instructions, even if its identity consistency or physical plausibility is high, its final score will still be significantly suppressed. Simultaneously, when an object's identity shifts or its movement violates physical laws, the overall score will also decrease. Finally, the video with the highest joint score exceeding a preset threshold is selected as the target demonstration sequence. Through this cross-dimensional coupled evaluation mechanism, the synergistic optimization of semantic understanding, target preservation, and physical executability is achieved, avoiding the misselection problem caused by mutual compensation among indicators in traditional independent scoring methods, and improving the reliability and executability of the target demonstration sequence.
[0034] S2. Construct cross-frame associations and recover the 3D structure of objects: Based on the initial observed image and the target demonstration sequence, a cross-frame anchor-query matching relationship is constructed. Specifically, the initial RGB-D image is used as the anchor, and each frame in the video sequence is used as the query. A two-dimensional pixel correspondence is established through cross-frame matching.
[0035] When establishing cross-frame pixel correspondences, the initial observed image is first used as the anchor frame. Each frame in the generated video is then matched frame-by-frame, and cross-frame correspondences are constructed by combining appearance features and geometric constraints. Specifically, visibility constraints based on depth information are used to define the matching region for pixels in the anchor frame. and candidate pixels within the matching area Extract their feature descriptors respectively and And use cosine similarity for matching evaluation: ; in, Indicates the similarity between two feature descriptors. Represents the vector dot product. and These represent the Euclidean norms of the corresponding feature vectors. Candidate points with similarity greater than a preset threshold are selected as the candidate matching point set. Furthermore, by introducing temporal consistency constraints, the matching results between adjacent frames are optimized to suppress matching drift and mismatches caused by the instability of the generated video. After obtaining stable two-dimensional pixel correspondences, the two-dimensional pixels are back-projected into three-dimensional space using depth information and camera intrinsic parameters from the initial observations to recover the corresponding three-dimensional point coordinates, thereby constructing a three-dimensional point set representation of the target object. Further, to improve the stability and scale consistency of 3D reconstruction, consistency screening and optimization are performed on the 3D points recovered across frames. By removing outliers and performing multi-frame fusion, the 3D point set maintains continuity and consistency in spatial structure, thus providing a reliable geometric basis for subsequent pose estimation. Compared to traditional methods based solely on single-frame or adjacent-frame matching, this method effectively reduces the accumulation of cross-frame matching errors in the time dimension by introducing anchor point constraints and temporal consistency optimization, improving the stability and accuracy of 3D structure recovery.
[0036] By constructing a three-dimensional implicit representation of an object, multi-frame visual information is mapped to a unified three-dimensional feature space and temporally aligned within this space to estimate the dynamic pose changes of the object. This can enhance the adaptability to complex deformation and occlusion scenes to some extent. However, it usually relies on multi-view data or pre-acquired information, which limits its application in single-view or unknown environments. In addition, this type of method has high computational complexity and is difficult to meet real-time requirements in practical applications.
[0037] S3. Perform model-free pose estimation and construct the object's pose trajectory: like Figure 2As shown, based on the aforementioned 3D point set representation, frame-by-frame pose estimation is performed on the video sequence to obtain the six-DOF pose information of the target object at continuous time scales. Specifically, a coarse-scale alignment process is first performed, which obtains the initial pose estimate of the target object in the current frame by performing global geometric matching on the 3D point set. The coarse-scale alignment adopts a matching strategy based on spatial distribution consistency, searching for the optimal pose transformation within a large range to improve robustness to initial pose deviations and scale uncertainties. Subsequently, a fine-grained alignment process is performed. Based on the coarse alignment result, the pose parameters are iteratively optimized by minimizing the reprojection error or inter-point distance error between the current frame observation point and the reference point set to improve registration accuracy and restore the fine-grained structural alignment relationship.
[0038] Furthermore, to address the issues of noise, jitter, and drift in frame-by-frame estimation of video sequences, a continuous optimization mechanism based on temporal consistency is introduced. Specifically, pose changes between adjacent frames are modeled as a continuous motion process. By constructing cross-frame constraints, the pose sequence is jointly optimized to ensure that pose changes remain smooth in the temporal dimension and conform to the laws of physical motion. Based on this, temporal smoothing and consistency optimization processing is performed on the pose sequence, including filtering translation and rotation parameters separately, and using an outlier detection mechanism to remove abrupt frames, thereby suppressing local anomalies caused by matching errors or instability in the generated video. Compared to traditional frame-by-frame independent estimation methods, this method, through a "coarse-to-fine" hierarchical registration and cross-frame joint optimization strategy, effectively reduces the accumulation of pose estimation errors in the temporal dimension, significantly improving the continuity, stability, and physical consistency of the trajectory.
[0039] By extracting sparse feature point trajectories from visual sequences, combining them with depth information to recover 3D points, and obtaining the six-degree-of-freedom pose of the object through outlier removal and pose solving methods, the computational complexity is reduced and robustness is improved to some extent. However, due to the reliance on sparse feature point matching, feature point loss or unstable matching is prone to occur when there is occlusion, violent movement, or viewpoint change, which affects the accuracy of pose estimation and trajectory continuity.
[0040] S4. Constructing a simulation task environment based on object pose trajectory: Based on initial visual observation information and the object's pose trajectory, a task execution scenario is constructed in a simulation environment. This simulation environment includes the scene's geometric structure, the initial state of the target object, and its physical properties, maintaining consistency with the real environment. The object's pose trajectory is introduced into the simulation environment as the task objective, enabling the robot's policy learning to drive the target object to evolve along the pose trajectory towards an optimization objective.
[0041] S5. Construct a reward function based on trajectory consistency and perform reinforcement learning training: like Figure 3 As shown, in the simulation environment, a reward function system with the target trajectory as the core constraint is constructed to guide the robot's policy learning.
[0042] The reward function includes, but is not limited to, proximity reward. Trajectory Consistency Reward Track progress rewards Etc. Among them, proximity reward Used to measure the distance relationship between the robot's end effector and the target object; trajectory consistency reward. Used to measure the deviation between the object's current pose and the target trajectory's pose at the corresponding moment; trajectory progress reward This is used to measure the extent to which an object progresses along a target trajectory, encouraging the strategy to complete the full motion process. The reward function is uniformly based on the object trajectory definition, rather than relying on manually taught trajectories or robot joint space information.
[0043] S6, Visual Strategy Construction and Cross-Domain Migration Execution: Based on the strategy trained in the simulation environment, a vision-driven robot execution strategy is constructed. Specifically, the state-action data in the simulation environment is converted into an image-action mapping relationship, and a visual strategy model is trained so that it can directly output control commands based on a single frame of visual input. The visual strategy is deployed to a real robot system, and a closed-loop control mechanism is constructed by combining real-time visual observations and online pose estimation results.
[0044] During execution, based on the difference between the real-time estimated object pose and the target trajectory, the position deviation and attitude deviation are calculated respectively. The position deviation is calculated by the difference between the target position in the target pose and the real-time estimated position. ; in This represents the target position at the time corresponding to the target trajectory. To estimate the current position of the object in real time, Let be the position deviation vector. Further, the position deviation magnitude can be expressed as: ; Attitude deviation is calculated as the angle difference between the target attitude and the real-time estimated attitude: ; in For the target attitude angle, To estimate attitude angles in real time, This represents the attitude deviation. If represented using three-dimensional Euler angles, the attitude deviation can be further expressed as: ; in: For roll angle, The pitch angle. The yaw angle is used. The calculated position and attitude deviation amplitudes are used for grading based on their trends. When the deviation is within a preset tolerance range, the current control strategy is maintained, with only minor adjustments to the output action to suppress error accumulation. When the deviation exceeds the first threshold and the change is continuously controllable, a local online adjustment mechanism is triggered. A feedback correction term for the trajectory error is introduced based on the current strategy output to incrementally compensate the robot's end effector motion, gradually returning the object to the target trajectory. When the deviation exceeds the second threshold or a sudden change occurs, it is determined that the current execution state has deviated from the stable trajectory, triggering a backtracking and replanning mechanism. The system state is reverted to the most recently stable historical node, and a new sequence of subsequent execution actions is generated based on that node.
[0045] During the online adjustment process, a control feedback model centered on trajectory deviation is constructed to decouple position and attitude errors. Furthermore, time continuity constraints are incorporated to smoothly transition correction actions, preventing control jitter or over-correction. Simultaneously, the error change rate is introduced as an auxiliary criterion to distinguish between gradual deviations and sudden disturbances, thereby adaptively selecting the appropriate correction strategy. Through this hierarchical adjustment and closed-loop optimization mechanism, the robot can stably recover from trajectory deviations in complex environments, significantly improving the robustness and accuracy of task execution.
[0046] Using object pose trajectory as a unified intermediate representation faces two key technical barriers in existing technologies: First, the data representation between different modules is inconsistent. Generative task demonstrations output pixel-level visual sequences, while robot control relies on continuous executable state-space representations, lacking a unified mapping relationship between the two. Second, model-free object perception struggles to obtain stable, continuous, and realistically scaled temporal poses under single-view conditions, leading to trajectory drift and inconsistencies, thus rendering it unreliable as a control basis. To address these issues, this embodiment constructs a unified trajectory parsing mechanism of "visual sequence—3D reconstruction—temporal pose optimization," transforming implicit motion information in generated videos into physically consistent six-DOF object pose trajectories, achieving a unified expression of cross-modal data. Building upon this, a further challenge arises: the coupling between generation, perception, policy learning, and execution becomes difficult. Specifically, the optimization objectives at each stage are inconsistent, and errors are difficult to propagate and correct. Traditional methods connect modules only through cascading, making it difficult to form a closed loop. To address this, this embodiment uses the object's pose trajectory as the sole shared constraint variable, simultaneously serving as the perception output, simulation learning target, and real execution feedback benchmark. This constructs a consistent optimization objective across stages and achieves online correction of trajectory deviations through a closed-loop feedback mechanism. This unifies the process from task generation to robot execution, avoiding simple additive integration and forming an end-to-end consistent technical solution. It improves the overall consistency and generalization capability of the system at the architectural level, while eliminating the need for manual teaching, object 3D models, and human posture detection, significantly reducing data acquisition costs and system deployment complexity.
[0047] This embodiment uses the six-degree-of-freedom pose trajectory of an object as an intermediate representation of the unified constraint signal and cross-morphological migration throughout the entire process, so as to achieve consistent constraints between the task objective, perception output and control process. Through policy learning in the simulation environment, the visual demonstration information is transformed into a control policy that the robot can execute, which effectively reduces the mapping error introduced by the difference between human and machine kinematics, thereby improving the trajectory tracking accuracy, execution stability and operation safety.
[0048] This embodiment employs a model-free six-DOF pose estimation method to achieve continuous tracking of the target object's motion process. The temporal pose trajectory of the unknown object can be extracted based on only a single frame of RGB-D initial observation. A closed-loop feedback adjustment mechanism is constructed in conjunction with the object's pose trajectory to effectively suppress the effects of tracking drift and environmental disturbances, thereby improving the system's robustness and adaptability in complex scenarios such as occlusion, viewpoint changes, and illumination changes.
[0049] This embodiment decouples the robot control strategy from the specific structure of the robot based on a unified object pose trajectory representation. This allows the control strategy learned in the simulation environment to be transferred to different robotic arm platforms and end effectors, demonstrating good versatility. At the same time, the complexity and computational overhead of the strategy learning process are reduced by the unified pose trajectory constraints, and stable and efficient execution performance can still be maintained in complex operation tasks.
[0050] This embodiment uses natural language commands and single-frame visual observations as inputs, and relies on the unified intermediate representation of object pose trajectory to realize an automated processing flow from task understanding to control execution. It can complete the robot's autonomous control tasks without the need for a large amount of prior data and manual annotation, and has good application prospects in unstructured environments such as home services, flexible manufacturing and smart warehousing.
[0051] Example 2: This embodiment provides a model-free robot control method based on object pose trajectory, including: The data acquisition module is configured to acquire initial visual observation information and task semantic information; The demonstration video determination module is configured to: generate candidate task demonstration videos based on the initial visual observation information and the task semantic information; and select target demonstration videos that describe the expected motion process of the target object from the candidate task demonstration videos. The execution strategy determination module is configured to: perform model-free pose estimation based on the initial visual observation information and the target demonstration video to obtain the six-degree-of-freedom pose trajectory of the target object in a continuous time scale; use the six-degree-of-freedom pose trajectory as a unified intermediate representation to train the robot strategy in a simulation environment and determine the execution strategy; The control module is configured to: control the robot to perform tasks based on the execution strategy and obtain the current pose of the object; and perform closed-loop adjustment of the robot's actions based on the deviation between the current pose of the object and the six-degree-of-freedom pose trajectory.
[0052] The working method of the system is the same as that of the model-free robot control method based on object pose trajectory in Embodiment 1, and will not be repeated here.
[0053] Example 3: This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the model-free robot control method based on object pose trajectory described in Embodiment 1.
[0054] Example 4: This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the steps of the model-free robot control method based on object pose trajectory described in Embodiment 1.
[0055] Example 5: This embodiment provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the model-free robot control method based on object pose trajectory described in Embodiment 1.
[0056] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.
Claims
1. A model-free robot control method based on object pose trajectory, characterized in that, include: Acquire initial visual observation information and task semantic information; Based on the initial visual observation information and the task semantic information, a candidate task demonstration video is generated; Target demonstration videos that describe the expected motion process of the target object are selected from the candidate task demonstration videos; Based on the initial visual observation information and the target demonstration video, model-free pose estimation is performed to obtain the six-degree-of-freedom pose trajectory of the target object on a continuous time scale. Using the six-degree-of-freedom pose trajectory as a unified intermediate representation, the robot strategy is trained in a simulation environment to determine the execution strategy; The robot is controlled to perform tasks based on the execution strategy and the current pose of the object is obtained; the robot's actions are adjusted in a closed loop based on the deviation between the current pose of the object and the six-degree-of-freedom pose trajectory.
2. The model-free robot control method based on object pose trajectory as described in claim 1, characterized in that, The model-free pose estimation includes: constructing point-level motion information of the object by predicting the two-dimensional point correspondence between the starting frame and the target frame, and recovering the three-dimensional point coordinates by combining the depth information; and estimating the six-degree-of-freedom pose trajectory of the object based on the three-dimensional point coordinates.
3. The model-free robot control method based on object pose trajectory as described in claim 2, characterized in that, The recovery of the three-dimensional point coordinates includes: using the initial observation image as the anchor frame, performing frame-by-frame matching on each frame of the generated video, constructing cross-frame correspondence by combining appearance features and geometric constraints; after obtaining a stable two-dimensional pixel correspondence, combining the depth information and camera intrinsic parameters in the initial observation, back-projecting the two-dimensional pixels to three-dimensional space, recovering the corresponding three-dimensional point coordinates, thereby constructing a three-dimensional point set representation of the target object.
4. The model-free robot control method based on object pose trajectory as described in claim 3, characterized in that, Based on the aforementioned 3D point set representation, the initial pose estimate of the target object in the current frame is first obtained by performing global geometric matching on the 3D point set; then, the pose parameters are iteratively optimized by minimizing the reprojection error between the observation point in the current frame and the reference point set.
5. The model-free robot control method based on object pose trajectory as described in claim 1, characterized in that, The initial visual observation information is an RGB-D image, and the task semantic information is a natural language instruction.
6. The model-free robot control method based on object pose trajectory as described in claim 1, characterized in that, The generation of the candidate task demonstration video includes: based on a temporal generation network, using initial visual observation information as scene content constraints and task semantic information as behavioral condition input, mapping language instructions to visual action priors through cross-modal conditional encoding, guiding the model to generate dynamic video sequences that conform to task semantics; in the process of generating dynamic video sequences, multiple sets of candidate video sequences are obtained by fixing the conditions of the initial frame and applying consistency constraints in the temporal dimension.
7. The model-free robot control method based on object pose trajectory as described in claim 1, characterized in that, The determination of the target demonstration video includes: for each candidate video in the candidate video sequence, using semantic consistency as a prior constraint to filter out sequences that do not meet the task objective, and then performing joint optimization and filtering based on object identity consistency and physical rationality.
8. The model-free robot control method based on object pose trajectory as described in claim 1, characterized in that, The closed-loop adjustment includes: during execution, determining the position deviation and attitude deviation based on the difference between the real-time estimated object pose and the target trajectory, and maintaining the current control strategy when the deviation is within a preset range.
9. The model-free robot control method based on object pose trajectory as described in claim 8, characterized in that, When the deviation exceeds the first threshold, a local online adjustment mechanism is triggered. Based on the current strategy output, a feedback correction term is introduced to incrementally compensate the robot's end effector motion, so that the object gradually returns to the target trajectory. When the deviation exceeds the second threshold, a rollback and replanning mechanism is triggered, which backtracks the state to the most recent historical node that met the stability conditions, and regenerates the subsequent execution action sequence based on the historical node.
10. A model-free robot control method based on object pose trajectory, characterized in that, include: The data acquisition module is configured to acquire initial visual observation information and task semantic information; The demonstration video determination module is configured to generate candidate task demonstration videos based on the initial visual observation information and the task semantic information. Target demonstration videos that describe the expected motion process of the target object are selected from the candidate task demonstration videos; The execution strategy determination module is configured to: perform model-free pose estimation based on the initial visual observation information and the target demonstration video to obtain the six-degree-of-freedom pose trajectory of the target object in a continuous time scale; Using the six-degree-of-freedom pose trajectory as a unified intermediate representation, the robot strategy is trained in a simulation environment to determine the execution strategy; The control module is configured to: control the robot to perform tasks based on the execution strategy and obtain the current pose of the object; The robot's movements are adjusted in a closed loop based on the deviation between the object's current pose and the six-degree-of-freedom pose trajectory.
Citation Information
Patent Citations
Human demonstration robot operation learning system and method combining visual sense and touch sense of object
CN121052281A
Industrial robot three-dimensional visual guidance AI accurate control method
CN122143060A