Visual servoing-based robot flow motion closed-loop control method and terminal equipment
Patent Information
- Application Number
- CN202610942681.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-06-29
AI Technical Summary
[0008]本发明的目的在于,提供一种基于视觉伺服的机器人流式动作闭环控制方法及终端设备,特别涉及一种结合高帧率视觉感知与动态动作序列规划的机器人控制方法及终端设备,更具体地说,涉及一种基于视觉伺服、利用Transformer模型预测抓取姿态,并通过流式动作块动态更新机制实现机器人闭环控制的方法及终端设备,从而解决现有机器人控制方法存在的缺乏对动态目标的在线适应能力、实时性、平滑性不足、对模型精度和计算资源要求过高的技术问题
[0033] (1) Achieve true dynamic closed-loop control and greatly improve the success rate of grasping moving targets: Real-time perception of the pose changes of target objects through high frame rate vision, and the adoption of a unique "streaming action block dynamic update" mechanism, which enables the robot's motion planning to adapt to target changes online and in real time, overcomes the fundamental defects of traditional open-loop execution and sluggish replanning methods, and brings core technical advantages for dealing with dynamic scenarios.
Smart Images

Figure CN122442693B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of robot control technology, and specifically discloses a robot flow motion closed-loop control method and terminal device based on visual servoing. Background Technology
[0002] In the field of robot grasping and manipulation, visual servoing is a key technology for achieving precise and adaptive control. Existing technologies mainly suffer from the following solutions and corresponding shortcomings:
[0003] 1. Open-loop "see-move" execution mode: The robot first acquires the pose of the target object through the vision system, and then plans and executes a fixed joint trajectory from the starting point to the ending point. The disadvantage of this mode is that it cannot handle dynamic changes such as displacement and rotation of the target object during the robot's movement, leading to grasping failure or a significant decrease in accuracy. It lacks online feedback and adjustment capabilities.
[0004] 2. Position-Based Visual Servoing (PBVS): This method estimates the target's pose in the robot's coordinate system in real time using a vision system and directly uses this pose as the controller's setpoint. Its drawback is that it is extremely sensitive to camera calibration accuracy; calibration errors directly translate into control errors. Furthermore, this method typically relies on traditional geometric feature matching, resulting in poor robustness to complex targets in unstructured environments or situations with occlusion.
[0005] 3. Replanning-based response strategy: When a change in target pose is detected to exceed a threshold, the current execution is interrupted, and global path planning is re-performed based on the robot's current state and the new target pose. Its disadvantages are that replanning involves large computational loads, introducing significant decision delays and motion pauses, resulting in choppy operation, low efficiency, and potentially falling into an oscillating state of "frequent interruptions and replanning" when the target changes continuously.
[0006] 4. Traditional Model Predictive Control (MPC): While capable of handling dynamic targets, it typically relies on accurate robot dynamics and environment models, and model mismatch can impact performance. Furthermore, its real-time optimization problem-solving is computationally burdensome, limiting its application in high-speed, high-frame-rate visual feedback scenarios.
[0007] In summary, the core problems with existing technologies are: either a lack of online adaptability to dynamic targets (such as open-loop execution), insufficient real-time performance and smoothness of the adaptation process (such as replanning), or excessively high requirements for model accuracy and computational resources. Therefore, there is an urgent need for a robot closed-loop control method that can integrate high frame rate visual perception, possess online generation and smooth switching capabilities for dynamic action sequences, and be relatively robust to model errors. Summary of the Invention
[0008] The purpose of this invention is to provide a robot streaming motion closed-loop control method and terminal device based on visual servoing. In particular, it relates to a robot control method and terminal device that combines high frame rate visual perception and dynamic motion sequence planning. More specifically, it relates to a method and terminal device based on visual servoing, using a Transformer model to predict grasping posture, and realizing robot closed-loop control through a streaming motion block dynamic update mechanism. This solves the technical problems of existing robot control methods, such as lack of online adaptability to dynamic targets, insufficient real-time performance and smoothness, and excessive requirements for model accuracy and computing resources.
[0009] The first aspect of the present invention provides a closed-loop control method for robot streaming motion based on vision servoing, comprising:
[0010] Step 1: Obtain the position of the target object in the camera coordinate system, and predict the grasping posture of the robot end effector based on the position, which is recorded as the initial grasping posture.
[0011] Step 2: Determine the first spatial trajectory from the robot's current joint angle to the joint angle corresponding to the initial grasping posture, and divide the first spatial trajectory into multiple consecutive action blocks;
[0012] Step 3: Determine the target grasping posture of the robot end effector when executing each action block in sequence, and determine whether the deviation between the target grasping posture and the initial grasping posture exceeds a threshold; if so, plan a second spatial trajectory from the joint angle corresponding to the current action block to the joint angle corresponding to the target grasping posture; the target grasping posture is the predicted grasping posture.
[0013] Step 4: Replace the remaining action blocks with the second spatial trajectory, and repeat steps 3 and 4 until the robot end effector completes the grasping.
[0014] Preferably, determining whether the deviation between the target grasping posture and the initial grasping posture exceeds a threshold specifically involves:
[0015] Determine whether the positional deviation or attitude angle deviation between the target grasping posture and the initial grasping posture exceeds the corresponding threshold.
[0016] Preferably, the threshold value of the positional deviation is 3mm - 6mm;
[0017] The threshold for the attitude angle deviation is 2°-4°.
[0018] Preferably, the first spatial trajectory is divided into multiple consecutive action blocks, specifically:
[0019] The first spatial trajectory is discretized into a first sequence containing multiple joint angle vectors;
[0020] The first sequence is evenly divided into multiple consecutive action blocks.
[0021] Preferably, a second spatial trajectory is planned from the joint angle corresponding to the current action block to the joint angle corresponding to the target grasping posture, specifically as follows:
[0022] The starting angle is obtained as the last joint angle that has been sent but not yet arrived within the action block.
[0023] Plan a second spatial trajectory from the starting angle to the joint angle corresponding to the target grasping posture.
[0024] Preferably, the remaining action blocks are replaced using the second spatial trajectory, specifically as follows:
[0025] The second spatial trajectory is discretized into a second sequence containing multiple joint angle vectors;
[0026] The second sequence is evenly divided into multiple consecutive action sub-blocks, the number of which corresponds to the number of remaining action blocks;
[0027] Replace the remaining action blocks with the action sub-blocks.
[0028] Preferably, the grasping posture of the robot's end effector is predicted based on the position, specifically as follows:
[0029] The position is input into the trained Transformer model to predict the grasping posture of the robot's end effector.
[0030] Preferably, the loss function of the Transformer model is the mean squared error function.
[0031] A second aspect of the present invention provides a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described visual servoing-based robot flow motion closed-loop control method.
[0032] The visual servoing-based robot flow motion closed-loop control method and terminal device of the present invention have the following advantages compared with the prior art:
[0033] (1) Achieve true dynamic closed-loop control and greatly improve the success rate of grasping moving targets: Real-time perception of the pose changes of target objects through high frame rate vision, and the adoption of a unique "streaming action block dynamic update" mechanism, which enables the robot's motion planning to adapt to target changes online and in real time, overcomes the fundamental defects of traditional open-loop execution and sluggish replanning methods, and brings core technical advantages for dealing with dynamic scenarios.
[0034] (2) Ensuring motion continuity and smoothness, and reducing system latency: The strategy of "updating subsequent unexecuted blocks during execution" is adopted to avoid motion interruptions and restarts caused by global replanning. The transition section design ensures continuous speed during trajectory switching, making the robot's motion always smooth and natural. The effect of this invention is to improve operational efficiency and the safety of human-machine collaboration.
[0035] (3) Reduce dependence on absolute positioning accuracy and enhance system robustness: The Transformer model used in this invention learns "relative pose mapping" (from object position to grasping posture) rather than absolute pose estimation that requires precise calibration. As long as the target moves relatively within the field of view, the model can predict a reasonable tracking and grasping posture, which has better robustness to camera calibration error and model absolute positioning error.
[0036] (4) Calculate load balancing to meet real-time requirements: Decompose long trajectories into action blocks for execution and restrict replanning to unexecuted local future trajectories, making the online computation amount controllable. Compared with the MPC method, which performs global planning in every control cycle, this invention significantly reduces computational complexity while ensuring dynamism, and is easier to run in real time on embedded controllers. Attached Figure Description
[0037] Figure 1 This is a flowchart of a robot flow motion closed-loop control method based on visual servoing, according to an embodiment of the present invention. Detailed Implementation
[0038] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0039] The first aspect of the present invention provides a closed-loop control method for robot streaming motion based on vision servoing, comprising:
[0040] Step 1: Obtain the position of the target object in the camera coordinate system, and predict the grasping posture of the robot's end effector based on the position, which is recorded as the initial grasping posture.
[0041] In this embodiment of the invention, a high frame rate vision sensor (such as an industrial camera) deployed in the robot's workspace is used to acquire images containing the target object in real time. Then, the position of the target object in the camera coordinates is determined based on the image (such as the center of a two-dimensional bounding box in the image or a three-dimensional coordinate point obtained by an auxiliary sensor).
[0042] Furthermore, the grasping posture of the robot's end effector is predicted based on the location, specifically:
[0043] The position is input into the trained Transformer model, which outputs the 6-DOF grasping posture of the robot's end effector relative to the target object. ,in Indicates the current moment. For location, The pose is represented by Euler angles.
[0044] The training process of the Transformer model in this embodiment of the invention is as follows:
[0045] Dataset Construction: This embodiment of the invention collects a large amount of demonstration data on humans performing grasping operations on specific types of objects. Each data point includes: the position of the target object in the camera coordinate system at the instant before grasping. or And the grasping posture of the robot end effector corresponding to the grasping action, recorded by a motion capture system or teach pendant. .
[0046] Model training: target location As the input sequence, the corresponding grasping posture As the output label, a Transformer model with an encoder-decoder structure is trained. This model learns a complex mapping from the object's position to the appropriate grasping pose, and its training loss function is... The mean squared error (MSE) function:
[0047] (1)
[0048] in, The total number of samples, For the first The actual grasping posture of each sample For the first Predicted grasping posture for each sample.
[0049] Step 2: Determine the first spatial trajectory from the robot's current joint angle to the joint angle corresponding to the initial grasping posture, and divide the first spatial trajectory into multiple continuous action blocks.
[0050] This embodiment of the invention assumes the initial grasping posture is The posture of grasping at all times Furthermore, using inverse kinematics (IK) and trajectory planning algorithms (such as cubic spline interpolation), a trajectory is planned from the current joint angle. To the initial grasping posture Corresponding joint angles The smooth joint spatial trajectory is denoted as the first spatial trajectory.
[0051] Furthermore, the first spatial trajectory is divided into multiple consecutive action blocks, specifically:
[0052] The first spatial trajectory is discretized into a first sequence containing multiple joint angle vectors; then the first sequence is uniformly divided into multiple consecutive action blocks.
[0053] For example, the first spatial trajectory is discretized into a form containing The first sequence of joint angle vectors Then Evenly divided into A series of action chunks: Each action block contains Joint angle points ( Can be Divisible by (divisible).
[0054] Step 3: Determine the target grasping posture of the robot's end effector when executing each action block sequentially, and determine whether the deviation between the target grasping posture and the initial grasping posture exceeds a threshold; if so, plan a second spatial trajectory from the joint angle corresponding to the current action block to the joint angle corresponding to the target grasping posture. The target grasping posture is the grasping posture predicted using the Transformer model in Step 1.
[0055] In this embodiment, the control system executes action blocks sequentially. During execution, the vision system continuously runs step 1, updating the target grasping pose at a high frame rate (e.g., 30Hz or higher). When it reaches step 1... When there is an action block (assuming this moment is...) ), monitoring mechanism judgment Target grasping posture at all times Compared to the initial grasping posture Does the deviation between them exceed the preset threshold? For example, the position deviation is greater than a first threshold or the attitude angle deviation is greater than a second threshold, where the first threshold is 3mm - 6mm, preferably 5mm; and the second threshold is 2° - 4°, preferably 3°. If the threshold is not exceeded, the original subsequent action block continues to be executed. If the threshold is exceeded, a streaming dynamic update is triggered: Record the first... The last joint angle point of the action block that has been sent but not yet arrived. (or use the actual joint angle fed back by the current motor) (This serves as the starting point for subsequent replanning) , and then with Starting from the angle, with the latest target attitude Corresponding joint angle Given the final state, perform local replanning to replan a joint space trajectory, denoted as the second space trajectory.
[0056] Step 4: Replace the remaining action blocks with the second spatial trajectory, and repeat steps 3 and 4 until the robot end effector completes the grasping.
[0057] In this embodiment of the invention, the second spatial trajectory is first discretized into a second sequence containing multiple joint angle vectors; then the second sequence is evenly divided into multiple continuous action sub-blocks, the number of which corresponds to the number of remaining action blocks; finally, the remaining action blocks are replaced with action sub-blocks.
[0058] For example, the second spatial trajectory is discretized into a form containing The second sequence of joint angle vectors; then the second sequence is evenly divided into A series of action sub-blocks The number of action sub-blocks corresponds to the number of remaining action blocks; finally, the remaining action blocks are replaced with action sub-blocks. .
[0059] To ensure a smooth transition from the old trajectory to the new trajectory, embodiments of the present invention, in A brief transition segment is inserted between the starting point of the new trajectory and the starting point. This transition segment is generated by linear interpolation or low-order polynomial fitting to ensure the continuity of joint position and velocity.
[0060] A second aspect of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described visual servoing-based robot flow motion closed-loop control method.
[0061] The control method of the present invention will be described in detail below with more specific embodiments.
[0062] Example 1
[0063] The visual servoing-based robot flow motion closed-loop control method of this invention includes the following steps:
[0064] Step 101: System Initialization. Start the high frame rate camera (2 megapixels, 60fps) and robot control system. Load the pre-trained grasping pose prediction Transformer model. Set the number of action blocks. First threshold Second threshold .
[0065] Step 102: Perception and Initial Prediction. The camera captures an image containing the target water glass. The detection algorithm provides the pixel coordinates of the center of the bottom of the water glass in the image. Input these coordinates into the Transformer model, and the model will output the initial grasping pose. For example, the position relative to the rim of the cup. (The end is vertically downward).
[0066] Step 103: Initial trajectory planning and grouping. Let the robot's current joint angle be... Solving through inverse kinematics Obtain the target joint angle Plan a trajectory that takes 5 seconds, discretizing it into 500 points ( Divided into 10 action blocks (). Each block contains 50 points ( This corresponds to a 0.5-second motion. The first action block begins execution. .
[0067] Step 104: Streaming Execution and Monitoring. During execution... During this period, the vision module continues to run. When execution reaches... At point 25 (i.e., the overall progress is...) The vision module detected that the cup had been pushed.
[0068] Step 105: Trigger dynamic update. The latest predicted pose is now displayed. Compared with the initial attitude In comparison, the location is The direction was offset by 20mm ( The closed-loop decision-maker triggers an update.
[0069] Step 106: Local replanning and replacement. Streaming action planner records. Joint angle of the last planning point (point 50) As a new starting point. and corresponding to New target joint angle Using the starting and ending points as references, a new trajectory with a remaining duration of approximately 3.5 seconds is planned, discretely divided into 350 points, and further subdivided into 7 new action blocks. To replace the original .
[0070] Step 107: Smooth Switch and Continue Execution. and A 0.1-second transition trajectory is inserted between the starting points. The robot controller completes... Then, the transition block and the new subsequent action block are executed seamlessly. This process continues until the moving cup is successfully grabbed.
[0071] The visual servoing-based robot flow motion closed-loop control method and terminal device of the present invention have the following beneficial effects:
[0072] (1) Achieve true dynamic closed-loop control and greatly improve the success rate of grasping moving targets: Real-time perception of the pose changes of target objects through high frame rate vision, and the adoption of a unique "streaming action block dynamic update" mechanism, which enables the robot's motion planning to adapt to target changes online and in real time, overcomes the fundamental defects of traditional open-loop execution and sluggish replanning methods, and brings core technical advantages for dealing with dynamic scenarios.
[0073] (2) Ensuring motion continuity and smoothness, and reducing system latency: The strategy of "updating subsequent unexecuted blocks during execution" is adopted to avoid motion interruptions and restarts caused by global replanning. The transition section design ensures continuous speed during trajectory switching, making the robot's motion always smooth and natural. The effect of this invention is to improve operational efficiency and the safety of human-machine collaboration.
[0074] (3) Reduce dependence on absolute positioning accuracy and enhance system robustness: The Transformer model used in this invention learns "relative pose mapping" (from object position to grasping posture) rather than absolute pose estimation that requires precise calibration. As long as the target moves relatively within the field of view, the model can predict a reasonable tracking and grasping posture, which has better robustness to camera calibration error and model absolute positioning error.
[0075] (4) Calculate load balancing to meet real-time requirements: Decompose long trajectories into action blocks for execution and restrict replanning to unexecuted local future trajectories, making the online computation amount controllable. Compared with the MPC method, which performs global planning in every control cycle, this invention significantly reduces computational complexity while ensuring dynamism, and is easier to run in real time on embedded controllers.
[0076] The above descriptions are merely a few embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any modifications or alterations made by those skilled in the art without departing from the scope of the technical solution of the present invention using the disclosed technical content are equivalent to equivalent implementation cases and fall within the scope of the technical solution.
Claims
1. A closed-loop control method for robot streaming motion based on visual servoing, characterized in that, include: Step 1: Obtain the position of the target object in the camera coordinate system, and predict the grasping posture of the robot end effector based on the position, which is recorded as the initial grasping posture. Step 2: Determine the first spatial trajectory from the robot's current joint angle to the joint angle corresponding to the initial grasping posture, and divide the first spatial trajectory into multiple consecutive action blocks; Step 3: Determine the target grasping posture of the robot's end effector when executing each action block sequentially, and determine whether the deviation between the target grasping posture and the initial grasping posture exceeds a threshold; if so, plan a second spatial trajectory from the joint angle corresponding to the current action block to the joint angle corresponding to the target grasping posture, specifically: obtain the last joint angle that has been sent but not yet arrived within the action block as the starting angle; plan the second spatial trajectory from the starting angle to the joint angle corresponding to the target grasping posture; the target grasping posture is the predicted grasping posture; Step 4: Replace the remaining action blocks with the second spatial trajectory, and repeat steps 3 and 4 until the robot end effector completes the grasping. The remaining action blocks are replaced using the second spatial trajectory, specifically as follows: The second spatial trajectory is discretized into a second sequence containing multiple joint angle vectors; The second sequence is evenly divided into multiple consecutive action sub-blocks, the number of which corresponds to the number of remaining action blocks; Replace the remaining action blocks with the action sub-blocks.
2. The visual servoing-based robot flow motion closed-loop control method according to claim 1, characterized in that, Determining whether the deviation between the target grasping posture and the initial grasping posture exceeds a threshold specifically involves: Determine whether the positional deviation or attitude angle deviation between the target grasping posture and the initial grasping posture exceeds the corresponding threshold.
3. The robot flow motion closed-loop control method based on visual servoing according to claim 2, characterized in that, The threshold for the positional deviation is 3mm-6mm; The threshold for the attitude angle deviation is 2°-4°.
4. The robot flow motion closed-loop control method based on visual servoing according to claim 1, characterized in that, The first spatial trajectory is divided into multiple consecutive action blocks, specifically: The first spatial trajectory is discretized into a first sequence containing multiple joint angle vectors; The first sequence is evenly divided into multiple consecutive action blocks.
5. The robot flow motion closed-loop control method based on visual servoing according to claim 1, characterized in that, The grasping posture of the robot's end effector is predicted based on the location, specifically as follows: The position is input into the trained Transformer model to predict the grasping posture of the robot's end effector.
6. The visual servoing-based robot flow motion closed-loop control method according to claim 5, characterized in that, The loss function of the Transformer model is the mean squared error function.
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the visual servo-based robot flow motion closed-loop control method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Target control method and device based on multi-modal information, equipment and medium
CN120816476A
Robot control method, robot and electronic equipment
CN121132643A