Collaborative robot end control method based on vision and related equipment thereof
Through the collaborative robot end control method of multimodal visual perception and real-time closed-loop feedback, the problem of low control accuracy in dynamic environments is solved, high-precision target pose estimation and path planning are achieved, and the robot's independent decision-making and execution capabilities in complex environments are improved.
Patent Information
- Application Number
- CN202510617743.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing collaborative robot end control methods lack depth perception and real-time integration of spatial scenes and object semantics in dynamic and complex environments, making it difficult to accurately estimate the target pose and obstacle distribution, resulting in low control accuracy.
Using multimodal visual perception, semantic three-dimensional modeling and real-time closed-loop feedback methods, the original visual data is dynamic anti-interference filtering and semantic fusion processing, and the three-dimensional semantic point cloud data is obtained, coordinate system alignment, path planning and inverse kinematics are performed, and closed-loop feedback optimization is performed by combining visual servo and impedance control.
It significantly improves the robot's independent decision-making and execution capabilities in complex and dynamic environments, realizes seamless mapping of visual data to mechanical execution, reduces impact and vibration during path tracking, and improves positioning accuracy and motion smoothness.
Smart Images

Figure CN120287303A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of device control, and in particular, to a vision-based end control method for a collaborative robot and related devices. Background Art
[0002] In various scenarios such as industrial production, medical assistance, and the service industry, collaborative robots need to operate in parallel with humans or other devices in a dynamic, complex, and unpredictable environment. Therefore, extremely high requirements are imposed on the real-time perception, precise planning, and closed-loop correction of the end position and posture.
[0003] The end control methods mainly include position-based open-loop control, force / torque-based force control or impedance control, and a hybrid control strategy of the two. Open-loop control relies on a preset trajectory, which is simple in structure but lacks environmental adaptability; pure force control optimizes the contact behavior through force sensor feedback, but it is difficult to balance spatial positioning accuracy; impedance control balances rigidity and flexibility with a virtual mass-damping-spring model, but has a low dependence on vision and spatial information and often requires a complex tactile or high-density sensing layout; the hybrid control compensates for the deficiencies of a single strategy to a certain extent, but increases the system coupling complexity and the difficulty of parameter tuning.
[0004] The existing end control methods lack in-depth perception and real-time fusion of the spatial scene and object semantics, and it is difficult to accurately estimate the target pose and obstacle distribution in a dynamic and complex working environment. Summary of the Invention
[0005] In view of this, the present application provides a vision-based end control method for a collaborative robot and related devices to solve the problem of low end control accuracy in a dynamic environment.
[0006] The first aspect of the present application provides a vision-based end control method for a collaborative robot, and the method includes: Performing multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original visual data to obtain three-dimensional semantic point cloud data; Performing coordinate system alignment processing on the three-dimensional semantic point cloud data according to a preset robot base coordinate system to obtain a set of target pose parameters in the robot base coordinate system; Performing dynamic path planning processing on the set of target pose parameters to obtain a joint space motion trajectory; Performing inverse kinematics solution and motor command synthesis processing on the joint space motion trajectory according to preset robot D-H model parameters to obtain a motor control command; Performing vision-based servo and impedance control closed-loop feedback optimization processing on the motor control command according to the state data of the end effector collected in real time to obtain a dynamic correction control command.
[0007] In an alternative embodiment, the multi-modal dynamic anti-interference filtering and semantic fusion processing of the collected original visual data to obtain three-dimensional semantic point cloud data includes: Performing adaptive bilateral filtering on the collected original visual data according to a preset spatial distance standard deviation, pixel value difference standard deviation, and normalization coefficient to obtain denoised visual data; Performing coordinate deviation alignment on the denoised visual data through a preset checkerboard calibration model to obtain aligned and optimized visual data; Performing object detection on the aligned and optimized visual data through a preset YOLOv7 model to obtain object bounding boxes and semantic labels; Performing depth information mapping on the aligned and optimized visual data according to a preset camera internal parameter to obtain spatial point cloud data, and performing data encapsulation on the spatial point cloud data, the object bounding boxes, and the semantic labels according to a preset data format to obtain the three-dimensional semantic point cloud data.
[0008] In an alternative embodiment, the target pose parameter set in the robot base coordinate system includes target coordinate data and object pose angle data. The coordinate system alignment of the three-dimensional semantic point cloud data according to the preset robot base coordinate system to obtain the target pose parameter set in the robot base coordinate system includes: Calculating the camera-robot coordinate system transformation matrix through the checkerboard calibration model according to the preset robot base coordinate system by the least squares method; Performing transformation on the object coordinate data in the three-dimensional semantic point cloud data according to the camera-robot coordinate system transformation matrix to obtain the target coordinate data in the robot base coordinate system; Performing principal component analysis on the target coordinate data to obtain the object pose angle data.
[0009] In an alternative embodiment, the dynamic path planning of the target pose parameter set to obtain a joint space motion trajectory includes: Constructing an occupancy grid map for the target pose parameter set to obtain an obstacle marked area; Performing global path search on the target pose parameter set according to a preset improved A* algorithm and the obstacle marked area to obtain optimal discrete path points; Performing B-spline curve fitting on the optimal discrete path points according to a preset acceleration constraint to obtain the joint space motion trajectory.
[0010] In an optional embodiment, the inverse kinematics solution and motor command synthesis processing of the joint space motion trajectory according to the preset robot D-H model parameters to obtain the motor control command include: Construct a kinematic model according to the preset robot D-H model parameters to obtain a forward kinematics formula; Perform pose mapping on the joint space motion trajectory according to the forward kinematics formula to obtain end-effector pose data; Construct an objective function according to the preset target end-effector pose data and the end-effector pose data, and use the Levenberg–Marquardt algorithm to perform inverse kinematics iterative solution on the objective function to obtain each joint angle; Convert each joint angle into a motor pulse signal according to the preset encoder pulse number and reduction ratio, and perform synthesis processing on the motor pulse signals according to the time sequence to obtain the motor control command.
[0011] In an optional embodiment, the end-effector state data includes real-time visual data and contact force data. The closed-loop feedback optimization processing of the motor control command based on vision and impedance control according to the real-time collected end-effector state data to obtain the dynamic correction control command includes: Perform target detection processing on the real-time visual data to obtain a set of target feature points; Calculate the feature point deviation according to the target end-effector pose data and the set of target feature points, and perform pseudo-inverse calculation of the Jacobian matrix on the feature point deviation to obtain a joint angle correction amount; Adjust the end-effector position deviation of the contact force data through a preset impedance control model to obtain a force control position correction amount; Perform superposition optimization processing on the motor control command according to the joint angle correction amount and the force control position correction amount to obtain the dynamic correction control command.
[0012] In an optional embodiment, the method further includes: Perform human pose tracking detection on the real-time visual data to obtain skeleton key points; Calculate the distance according to the received key points on the robot surface and the skeleton key points to obtain the minimum distance; Compare the minimum distance with a preset distance threshold to obtain a corresponding emergency response signal.
[0013] The second aspect of the present application provides a vision-based collaborative robot end control device, and the device includes: A data fusion module, which is used to perform multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original visual data to obtain three-dimensional semantic point cloud data; A coordinate optimization module, which is used to perform coordinate system alignment processing on the three-dimensional semantic point cloud data according to a preset robot base coordinate system to obtain a set of target pose parameters in the robot base coordinate system; A path planning module, which is used to perform dynamic path planning processing on the set of target pose parameters to obtain a joint space motion trajectory; An instruction synthesis module, which is used to perform inverse kinematics solution and motor instruction synthesis processing on the joint space motion trajectory according to preset robot D-H model parameters to obtain a motor control instruction; A feedback optimization module, which is used to perform vision-based servo and impedance control closed-loop feedback optimization processing on the motor control instruction according to the real-time collected end effector state data to obtain a dynamically corrected control instruction.
[0014] A third aspect of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned vision-based collaborative robot end control method are implemented.
[0015] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned vision-based collaborative robot end control method are implemented.
[0016] In summary, the present application at least includes the following beneficial technical effects: 1. By organically integrating multi-modal visual perception, semantic three-dimensional modeling, real-time closed-loop feedback, and efficient path planning, the autonomous decision-making and execution capabilities of the robot in complex and dynamic environments are significantly improved.
[0017] 2. By combining coordinate system alignment with principal component analysis, seamless mapping from visual data to mechanical execution is achieved, avoiding the cumulative errors caused by traditional calibration drift and multi-sensor fusion.
[0018] 3. By combining improved global path search with curve fitting with acceleration constraints, not only the global optimality of the trajectory is guaranteed, but also the smoothness of the output and the dynamic response speed are taken into account, significantly reducing the impact and vibration during path tracking. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of a vision-based collaborative robot end control method provided by an embodiment of the present application; Figure 2 It is a functional module diagram of a vision-based collaborative robot end control device provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0022] As Figure 1 shown, it is a flowchart of a vision-based collaborative robot end control method provided by an embodiment of the present application. The vision-based collaborative robot end control method provided by an embodiment of the present application includes the following steps.
[0023] Step S1: Perform multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original visual data to obtain three-dimensional semantic point cloud data.
[0024] In an industrial environment, RGB images, depth maps, and infrared images collected by a camera are often affected by interference such as dust, sudden changes in light, and mechanical vibrations. The embodiment of the present application uses adaptive bilateral filtering to not only suppress noise but also retain the edge features of objects. First, read the collected original visual data, where the original visual data includes RGB image I rgb (x), depth map I depth (x), and infrared image I ir (x). And filter each image type using the following formulas respectively: Among them, is the image coordinate point. is the neighborhood centered on x. is the preset standard deviation of the spatial distance, which is used to control the size of the filtering radius. is the preset standard deviation of the pixel value difference, which is used to control the similarity degree of grayscale (or depth, infrared value). is the preset normalization coefficient, which is equal to the sum of all weights and ensures that the range of pixel values remains unchanged after filtering. is the original image value at the neighborhood point y. is the new value of the filtered point x.
[0025] Spatial weighting makes the pixels farther from the center have less influence on the result, avoiding information aliasing when crossing the object edge; Intensity weighting makes the weight smaller when the difference in grayscale (or depth, infrared) is larger, so as to retain the edge. Eliminate environmental interferences such as uneven illumination and dust interference through adaptive bilateral filtering to improve the signal-to-noise ratio of the image. After filtering, the denoised visual data of the three-channel images is obtained .
[0026] Due to the slight deviations in the physical structure and field of view angle of the RGB, depth, and infrared sensors, the images collected by them must be spatially aligned in order to fuse the three-channel information under the same pixel coordinates. Specifically, a checkerboard calibration board with a known size is fixed on the central plane of the robot's end vision. The checkerboard images are taken separately using the RGB, depth, and infrared sensors at multiple different perspectives (at least 10), and the internal and external parameters are extracted. Calculate the internal parameter matrix K (i.e., focal length, principal point) and distortion coefficient of each camera through calibration algorithms such as OpenCV. Thus, calculate the external parameter matrix between the cameras to describe the coordinate deviation from camera 1 to camera 2. For the filtered Perform mapping interpolation according to inverse distortion + external parameter transformation to obtain the alignment results of the depth map and infrared map in the RGB coordinate system. Through the sub-pixel corner point detection of the grid points of the checkerboard, the alignment accuracy can be guaranteed to reach the 0.1 pixel level. External parameter calibration allows the accurate mapping of depth and infrared data to the RGB pixel plane, facilitating subsequent pixel-level data fusion. The aligned and optimized visual data is denoted as .
[0027] Quickly and accurately detect the boundaries and categories of different objects on the complete three-channel aligned and optimized visual data, and label the semantic point cloud. Specifically, load the pre-trained or fine-tuned YOLOv7 model, and at the same time scale and normalize the aligned RGB image according to the input size of the model. The YOLOv7 model performs forward inference layer by layer, and outputs a set of candidate boxes and their classification confidence 。Next, non-maximum suppression is adopted to remove redundant detection boxes with high overlap and low confidence. Each remaining box is attached with a class label (e.g., "Workpiece A", "Obstacle B"), as well as the corresponding pixel-level mask. Detecting each object through the YOLOv7 model provides a two-dimensional bounding box and class, which are projected and semantically labeled in the three-dimensional point cloud. The obtained detection results are denoted as { , }.
[0028] Finally, the two-dimensional pixel coordinates and depth values are converted into three-dimensional space coordinates, and each three-dimensional point is attached with color and semantic labels to achieve a three-dimensional semantic point cloud. Specifically, read the aligned depth map and RGB image . Using the known camera intrinsic parameters to map the pixel coordinates and the depth map into three-dimensional points in the camera coordinate system. Among them, . Attach color information to each three-dimensional point . According to the previously obtained two-dimensional bounding box { } and semantic label { }, the three-dimensional points falling within the same box are uniformly labeled as this class. And encapsulate the spectrum in the preset PLY format and output the three-dimensional semantic point cloud for subsequent path planning operations.
[0029] Step S2: Perform coordinate system alignment processing on the three-dimensional semantic point cloud data according to the preset robot base coordinate system to obtain the target pose parameter set in the robot base coordinate system.
[0030] The three-dimensional semantic point cloud data is initially represented based on the camera coordinate system. In actual operations, the motion instructions and path planning of the collaborative robot are all based on the robot base coordinate system. Therefore, it is necessary to accurately transform the point cloud data from the camera coordinate system to the robot base coordinate system. To achieve this goal, it is necessary to first determine the accurate mapping relationship between the two coordinate systems, that is, solve the coordinate transformation matrix from the camera to the robot base coordinate system. Specifically, prepare a checkerboard calibration board with known dimensions (e.g., each grid side length is 30mm), and place the calibration board in the operable working area of the robot. The calibration board needs to be firmly fixed to ensure no displacement. Measure the position information { } of the calibration board in the robot base coordinate system through the sensor system on the robot end effector or base, and at the same time measure the point cloud coordinates { }。At least 15 or more corresponding points are collected to ensure the stability and accuracy of the solution of the transformation matrix. Thus, the coordinates of two sets of points are input into the checkerboard calibration model, and the checkerboard calibration model, based on the least-squares fitting principle in Euclidean space, solves the transformation matrix (R, T) from the camera to the robot base coordinate system (i.e., the rotation matrix R and the translation vector T) according to these two sets of point coordinates. Among them, the rotation matrix R is a 3×3 rotation matrix, which is used to represent the rotation relationship between coordinate systems. The translation vector T is a 3×1 translation vector, which is used to represent the translational offset between coordinate systems.
[0031] Transform each three-dimensional point (i.e., point cloud) in the camera coordinate system to the robot base coordinate system to ensure that the robot can accurately identify the real-world position of the object when performing tasks such as grasping and obstacle avoidance. Specifically, for each point (x, y, z) in the three-dimensional semantic point cloud data, rotation and translation are performed according to the obtained transformation matrix (R, T), so as to accurately map each point from the position in the camera view to the reference system of the robot. Only in this way can the robot perform refined actions according to its own coordinates. All the transformed three-dimensional points, together with their respective semantic label information (i.e., class name, bounding box attribution), are saved to obtain the target coordinate data.
[0032] Obtain the main orientation of the object in three-dimensional space (for example, which direction the long side points to), which is crucial for the attitude control of the robot end effector. For example, when grasping a slender object, the grasping attitude must be consistent with the main axis direction of the object to ensure stable grasping. By extracting all the three-dimensional points corresponding to each semantic classification object (i.e., the set of points belonging to the same object), principal component analysis is applied to these three-dimensional point sets to calculate the main direction vector of the object. Specifically, the average position of all points in the set of points of the same object is calculated to obtain the centroid of the object. Then, the deviation vector of each point in the point set is calculated according to the centroid to construct the deviation matrix D of the object. Next, the covariance matrix of the object is calculated according to the deviation matrix D and the number of point sets to obtain the covariance matrix C of the object. Furthermore, the eigenvalue decomposition of the covariance matrix C is performed to obtain the eigenvectors and eigenvalues. Finally, the eigenvector corresponding to the largest eigenvalue is taken as the main axis direction vector of the object. By finding the direction with the largest extensibility (i.e., the main axis), the robot can determine the orientation of the object and adjust the attitude angle of the end effector accordingly.
[0033] Further, according to the main axis direction vector and in combination with the definition of the world coordinate system (i.e., the X-axis points forward, the Y-axis points left, and the Z-axis points upward), the rotation angle of the object (i.e., the object pose angle data) is calculated through trigonometric functions (such as the arctangent function arctan2). This object pose angle data, together with the corresponding object center position coordinates, forms the complete target pose parameter set (x, y, z, φ, θ, ψ) of the object, representing three-dimensional position and rotation angles around three axes respectively.
[0034] Step S3: Perform dynamic path planning processing on the target pose parameter set to obtain a joint space motion trajectory.
[0035] To convert complex point clouds and semantic information in three-dimensional space into a discrete representation suitable for path search, an occupancy grid map needs to be constructed. This occupancy grid map divides the workspace into three-dimensional grid cells of equal size (i.e., voxels) and marks whether each voxel is occupied by an obstacle, providing discrete nodes and obstacle boundaries for subsequent graph search. Specifically, according to the working radius and task range of the robot, three-dimensional boundaries are set. According to the preset grid resolution r (e.g., 10 mm), which serves as the side length of each voxel. The bounding box is divided along the three-dimensional coordinate axes according to the resolution to obtain Nx × Ny × Nz voxels. Initialize the marking value s for each voxel. i,j,k = 0, indicating "idle". Traverse all points (x, y, z) in the three-dimensional semantic point cloud and calculate the corresponding grid indices i, j, k for the point. Among them, . Set the corresponding voxel s i,j,k = 1 and mark it as "occupied". If the same voxel has already been marked, there is no need to recalculate. Further, the set of all voxels with s i,j,k = 1 is the obstacle marking area.
[0036] On the occupancy grid map, based on the voxel where the robot is currently located (starting point) and the target voxel (ending point), use the improved A* algorithm to search for a discrete path that is the shortest in distance and as far away from obstacles as possible. The improvement lies in combining the repulsive term of the artificial potential field (i.e., introducing a repulsive weight coefficient ), preventing the path from being too close to obstacles or getting stuck in local optima. Specifically, map the current end pose and the target pose to the grid indices , . For any node n to be expanded (on the voxel (i, j, k)), calculate the cost value by integrating the obstacle marking area information through the following formula: Among them, is the actual cost from the starting point to node n, usually taking the cumulative Euclidean distance. is a heuristic function that takes the straight-line distance from node n to the end point. is the repulsive force weight coefficient, which is used to adjust the degree of staying away from obstacles. { } is the set of the central coordinates of all voxels marked as obstacles. is the Euclidean distance from the center of the current node to the center of the obstacle node. By and it is ensured that the obtained path is the shortest. At the same time, by applying an inverse repulsive force to the obstacle points, the path is inhibited from sticking to the obstacles, thus enhancing safety.
[0037] The specific implementation process of the improved A* algorithm is as follows: First, add the starting node to the open list (Open) and keep the closed list (Closed) empty. In each iteration, a node with the minimum total cost is selected as the current node from the open list, and then this node is removed from the open list and added to the closed list. Subsequently, starting from the current node, all its adjacent nodes are enumerated. For each adjacent node, if it is already in the closed list or marked as an obstacle, the processing is skipped; otherwise, the actual cost from the starting point to this adjacent node and the heuristic estimated cost are used to recalculate the total cost, and at the same time, the inverse repulsive force term applied to all obstacle points is added to update the total cost . If this adjacent node has not entered the open list before, or the newly calculated actual cost is smaller than the previously recorded one, the parent node of this adjacent node is set as the current node, and its and values are updated and then it is added to the open list. After that, the algorithm repeats the above process: each time, the node with the minimum total cost is extracted from the open list as the new current node until the end point is expanded or the open list is empty. Once the end point is moved into the closed list, an optimal path containing discrete grid points can be obtained by tracing the parent node pointers of each node from the end point back to the starting point step by step.
[0038] After searching with the improved A* algorithm, a series of discrete grid points { } are obtained by tracing back from the end point to the starting point through the parent node pointers as the optimal discrete path points.
[0039] Although the discrete grid path points indicate a safe route, they contain turns and jumps, and the robot cannot execute them directly. It is necessary to smooth the discrete path to generate a continuous and differentiable joint space trajectory and ensure that the joint acceleration does not exceed the safety threshold. Specifically, the optimal discrete path points { } are converted into actual space coordinates { }, combined with the target attitude angle Form a complete pose sequence. At the same time, set M path nodes P k (including position and orientation), and according to the set B-spline order p and the appropriate knot vector { }, and solve the control points { } to construct a B-spline curve (i.e., the fitting trajectory function). Calculate the fitting trajectory function of the B-spline curve according to the Cox-deBoor recurrence formula, so as to fit the optimal discrete path points into a smooth curve without intersection, continuous curvature and no warping at the control points, as the joint space motion trajectory. Further, in order to meet the requirement that the acceleration of any joint satisfies the corresponding acceleration constraint, perform a second-order derivative operation on the B-spline fitting curve to obtain the acceleration vector. Specifically, stretch the time parameterization (increase the total trajectory duration) to reduce the acceleration peak. And add control points or increase the spline order at the high curvature points of the B-spline fitting curve to smooth the curvature change. Repeat the above fitting and verification until all accelerations satisfy the acceleration constraint.
[0040] Finally, obtain the joint space motion trajectory Q(t) in the form of a function that satisfies the position accuracy, speed and acceleration constraints, where each component q i (t) corresponds to the law of change of the angle of the i-th joint over time. The joint space motion trajectory can be discretized into sequential control instructions, which can be directly used as motor instruction inputs.
[0041] Step S4: Perform inverse kinematics solution and motor instruction synthesis processing on the joint space motion trajectory according to the preset robot D-H model parameters to obtain motor control instructions.
[0042] It should be understood that the D-H model parameters can establish a mathematical mapping between the geometric relationship of each link of the robot and the joint variables, and are used to calculate the spatial pose of the end effector at any joint angle. The preset robot D-H model parameters are stored in the form of a table, where each joint i ∈ {1,..., 6} corresponds to the following four fixed parameters: α j-1 : Link length, representing the distance from the i-th joint axis to the (i + 1)-th joint axis along the x i-1 axis. a j-1 : Link twist angle, representing the rotation angle of the x i-1 axis of the i-th link coordinate system to the x i axis around the x i-1 axis.
[0043] d j : Link offset, representing the offset from the i-th joint axis to the previous link along the z i-1 axis.
[0044] θ j: Joint variable, rotation angle about the z i-1 axis, which is a "movable" quantity.
[0045] Furthermore, it is necessary to define the homogeneous transformation matrix between adjacent links. The transformation between each pair of adjacent coordinate systems is represented by the homogeneous matrix as shown below: Among them, is the j-th joint angle (unknown quantity). α j-1 , a j-1 , d j are known quantities in the D-H parameters. The first three columns and the first three rows of the homogeneous matrix are responsible for rotation, the fourth column and the first three rows are responsible for translation, and the last row is used for homogeneous coordinates.
[0046] Furthermore, by multiplying the joint transformations in sequence, the forward kinematics formula for the homogeneous transformation from the base coordinate system to the end coordinate system is obtained: , where is the joint angle vector. The D-H parameterization method is a standard modeling method for industrial robots, which can separate complex spatial geometry into fixed parameters and joint variables. Through matrix multiplication, the three-dimensional position and orientation of the end effector can be obtained at once, without calculating axis by axis.
[0047] At the same time, the joint space trajectory gives a series of discrete time points ( = 0, ……, K) of the joint angle vector . For each time point substitute into the above forward kinematics formula to obtain . Furthermore, from the homogeneous matrix , the first three rows of the first 3×1 columns are the position coordinates [x, y, z] T , and the attitude angles (φ, θ, ψ) are extracted from the first three rows and the last three columns (3×3 submatrix) by methods such as Euler angles or quaternions. Using matrix multiplication to quickly calculate the end pose at each moment in batches to obtain the end pose data X current (t k ) = [x, y, z, φ, θ, ψ].
[0048] For each time point t k , combining the known target pose X desired (t k ) and the end pose X current (q(t k )), the following non-linear least squares objective function is constructed: Among them, is the end - pose vector obtained by forward kinematics calculation. is the target pose given by the path planning or vision module. is a preset regularization coefficient used to control the smoothness of joint solutions. is the joint angle at the previous moment, used to add continuity constraints to the solution.
[0049] In the embodiment of this application, the Levenberg–Marquardt algorithm is used to perform Gauss - Newton iteration to solve the objective function. After obtaining the initial joint angles from path planning or visual feedback the iteration count N is set to zero and an initial damping factor is selected. It should be understood that the objective function can be split into pose error residuals and regularization residuals Concatenating the pose error residuals and the regularization residuals can obtain a 12 - dimensional residual vector . At the same time, the Jacobian matrix J used in the solution process is the derivative of the residual vector with respect to the joint vector. By performing numerical differentiation or analytical differentiation on each output component in the robot forward kinematics formula with respect to each joint angle , to obtain . And construct a diagonal matrix , so as to and are concatenated vertically to obtain the Jacobian matrix J. The Jacobian matrix quantifies the influence of small changes in joint angles on the end - pose and regularization residuals. It should be understood that in the traditional Gauss - Newton method, in each iteration, by linearizing the residuals, the non - linear least - squares problem is converted into a linear equation. The traditional Gauss - Newton method converges quickly, but it is prone to divergence when the Jacobian matrix is close to singular or far from the initial value. In the embodiment of this application, on the basis of the traditional Gauss - Newton method, a damping term is added. The modified Levenberg–Marquardt algorithm used in the embodiment of this application can be represented by the following formula: where is the normal equation matrix of Gauss - Newton, reflecting second - order approximation information. λ is the damping coefficient, a scalar, used to balance Gauss - Newton and the steepest descent. Taking the diagonal elements of the normal equation matrix to form a diagonal matrix ensures reasonable values for the damping term. After adding the damping term, when is large, is approximately a diagonal matrix, and gradient descent is performed on to ensure stable convergence; when is very small, it degenerates into Gauss - Newton, and the convergence speed is faster.
[0050] The iterative update process is as follows: Before the iterative operation, it is necessary to obtain the initial joint angles and initialize and set the initial damping coefficient . The initial joint angles usually take the trajectory approximation value obtained in the previous step or the solution at the previous moment . The initial damping coefficient is set to be relatively large (for example, from 1e-2 to 1e-1) to ensure the stability of the initial iteration.
[0051] During the iterative operation, the objective function is disassembled and vector splicing calculations are performed in the above manner to obtain the current residual vector . At the same time, the current Jacobian matrix is calculated in the above manner . Furthermore, the increment is obtained according to the above modified Levenberg–Marquardt algorithm using an efficient linear algebra solution method (for example, Cholesky decomposition) . Thus, the increment is superimposed with the current joint angles to obtain , and then the new objective function value is calculated . If that is, the error decreases, accept the update . And through (and > 1), reduce the damping to approach the Gauss-Newton direction and accelerate the convergence. Otherwise, if the error does not decrease, reject the update and retain . Increase the damping , strengthen the gradient descent characteristic, and stabilize the convergence. By dynamically adjusting the damping coefficient, the algorithm mode can be adaptively switched to ensure correct convergence in both steep and flat regions of the curve.
[0052] After the iterative operation, if < (for example,[[]]ID=51]] ) or < (for example,[[]]ID=57]] ), it is considered that the joint angles have converged to the optimal solution, and output . Through the dual determination, both the stability of the solution is considered and it is ensured that the objective function obtains a sufficiently small error to meet the actual accuracy requirements.
[0053] Convert the joint angle solution into pulse signals recognizable by the actual execution driver, and assemble them into a complete control instruction sequence according to the time sequence. It should be understood that each joint motor consists of an encoder and a reduction mechanism, so the encoder pulses per revolution N need to be added to the motor pulse signalencoder , and the reduction ratio R gear . After obtaining the joint angles, the joint angles, encoder pulse counts, and reduction ratios are synthesized into signals through the following formula; where, is the angle of the -th joint at time . Divide by to convert radians to motor shaft revolutions.
[0054] Specifically, arrange the K of each joint in ascending order of time points t0, ……, t . For all joint pulses at the same time form a pulse vector , ……, . Pack the sequence into a data frame recognizable by the controller, usually containing fields such as timestamp, joint number, pulse value, checksum, etc., so as to form a continuous pulse sequence (i.e., motor control instruction) sent to the driver in time sequence.
[0055] Step S5: Based on the real-time collected end effector state data, perform vision-based servo and impedance control closed-loop feedback optimization processing on the motor control instruction to obtain a dynamically corrected control instruction.
[0056] Precisely extract the "target feature points" (such as workpiece edge corner points, cylindrical center line end points, etc.) that can be used for position deviation calculation from the real-time images collected by the camera, providing a basis for subsequent fine positioning and closed-loop control. First, apply grayscale conversion and Gaussian blur to each frame of real-time RGB image to remove noise, and use adaptive thresholding or Canny operator to extract edges to form a binary edge map. Then, input the edge map into a preset corner detection algorithm (e.g., Harris) and set a threshold. By calculating the response R of each pixel, retain the points with local maximum and greater than the preset threshold to obtain a preliminary corner set. Thus, associate the preliminary corner set with the bounding box and semantic labels, and only retain the corner points located within the current grasping target box to form a target feature point set.
[0057] Further, map the feature point deviation in the two-dimensional image or the three-dimensional point cloud deviation into the joint angle space to obtain the correction amount for driving the joint fine-tuning. For the real-time acquired image (i.e., real-time visual data), use the known three-dimensional pose of the target end and the current three-dimensional pose of the end. According to the camera internal parameters and distortion correction, back-project the pixel-level feature points into three-dimensional coordinates. Thus, obtain the corresponding real target feature three-dimensional points based on the three-dimensional coordinates, and the three-dimensional points measured after the current end hand-eye calibration. Calculate the difference between the real target feature three-dimensional points and the measured three-dimensional points to obtain the position error vector of each feature point. And construct the Jacobian matrix of the end position with respect to the joint angle according to the position error vector, and then use the Moore–Penrose algorithm to calculate the pseudo-inverse of the Jacobian matrix to obtain the joint angle correction amount 。
[0058] Meanwhile, in contact operations (such as grinding and assembly), perform compliant adjustment on the end position according to the real-time measured contact force to ensure operation safety and quality. Receive the contact force data output by the six-dimensional force / torque sensor at the end ,and use the linear impedance control model to input the force at the sensor sampling frequency into the second-order differential equation in the model for discretization processing (such as the forward Euler method). And obtain the steady-state correction amount through repeated iteration (i.e., the force control position correction amount). The impedance control model can be expressed by the following formula: where is the virtual mass matrix; is the damping matrix; is the stiffness matrix; is the incremental correction of the end position In order to combine the joint correction amount obtained by visual servoing with the force control position correction amount obtained by impedance control to form the final motor pulse adjustment. Among them, for the force control position correction amount
[0059] map it into the joint angle correction amount In a collaborative environment, it is necessary to continuously obtain the spatial position and posture of the operator. Skeletal key points (such as the positions of the head, torso, hand, and foot joints) can accurately describe the three-dimensional positions of various parts of the human body, providing basic data for distance calculation. Specifically, a binocular RGB-D camera is used to synchronously obtain the color image I rgb (u, v) and the aligned depth map I depth (u, v). The color image is undistorted using the internal parameters and distortion coefficients obtained from camera calibration, and converted to a grayscale image to reduce the computational load. The preprocessed color image is input into a pre-trained human pose estimation algorithm (e.g., OpenPose), and the algorithm outputs a set of two-dimensional key point coordinates {(u i , v i )|i = 1……N}, where i represents different human joint points (such as the top of the head, shoulders, elbows, wrists, hips, knees, ankles, etc.), and each point is accompanied by a confidence level c i . Thus, the human joint points are screened according to the confidence level c i , and only the key points with a confidence level c i higher than the threshold (e.g., 0.5) are retained to reduce false alarms. Further, the depth value d i , v i ) at each key point pixel position (u i = I depth (u i , v i ) is read using the aligned depth map. The pixel coordinates and depth are mapped to three-dimensional points (X i , Y i , Z i ) in the camera coordinate system according to the camera internal parameters. All the three-dimensional points are output as the human skeletal key point set for subsequent distance calculation.
[0060] The distances between the human key points and the discrete key point set on the surface of the robot itself are compared to find the closest human-robot point pair, which serves as the trigger basis for safety control. According to the robot CAD model or kinematic model, a set of key points on the robot surface (e.g., the center of the end effector, several representative points at the joint connections, the fuselage shell, etc.) are preset, and their coordinates have been determined in the robot base coordinate system. Further, the human skeletal key point set is mapped into the robot base coordinate system through the coordinate transformation matrix in the method of step S2. Thus, the three-dimensional distances between each mapped human skeletal key point and the robot surface key points are calculated one by one, and the minimum distance dmin among all the distance values is screened out.
[0061] Finally, according to the minimum distance and the predefined safety distance threshold, different emergency response signals (e.g., normal, deceleration, emergency stop) are triggered at different levels to ensure the safety of human-robot interaction under various proximity levels.
[0062] Exemplarily, two key distance thresholds are set in the embodiments of the present application, namely Dslow = 0.5 m (deceleration trigger distance) and Dstop = 0.3 m (emergency stop trigger distance). If dmin ≥ Dslow, the system maintains the "normal" mode and does not adjust the speed; if dmin ∈ [Dstop, Dslow), the system enters the "deceleration" mode and reduces the robot's moving speed to a preset ratio (e.g., 10%); if dmin < Dstop, the system enters the "emergency stop" mode, immediately stops the machine and commands the robot to retreat to a pre-stored safe point. At the same time, for the convenience of traceability inspection, a safety status flag Ssafe is generated according to the judgment result, and the safety status flag Ssafe and the corresponding trigger time tag are stored in a preset safety log together.
[0063] The present application is applied to the field of device control technology. By performing multi-modal dynamic anti-interference filtering and semantic fusion on the original visual data, three-dimensional semantic point cloud data is obtained. According to the robot base coordinate system, coordinate alignment is performed on the three-dimensional semantic point cloud data to obtain a set of target pose parameters. Dynamic path planning is performed on the set of target pose parameters to obtain a joint space motion trajectory. According to the robot D-H model parameters, inverse kinematics solution and motor command synthesis are performed on the joint space motion trajectory to obtain a motor control command. According to the end effector state data, closed-loop feedback optimization is performed on the motor control command to obtain a dynamically corrected control command. Through the multi-level processing and closed-loop feedback of visual information, the present application effectively improves the positioning accuracy, motion smoothness and interaction safety of the robot in a dynamic and complex environment.
[0064] As Figure 2 shown, it is a functional module diagram of a vision-based collaborative robot end control device provided by an embodiment of the present application.
[0065] In some embodiments, the vision-based collaborative robot end control device 2 may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the vision-based collaborative robot end control device 2 may be stored in the memory of the server and executed by at least one processor to execute (see Figure 1 the description) the functions of the vision-based collaborative robot end control method.
[0066] In this embodiment, the vision-based collaborative robot end control device 2 can be divided into multiple functional modules according to its executed functions. The functional modules may include: a data fusion module 21, a coordinate optimization module 22, a path planning module 23, an instruction synthesis module 24, a feedback optimization module 25, and a dynamic monitoring module 26. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0067] The data fusion module 21 is configured to perform multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original vision data to obtain three-dimensional semantic point cloud data.
[0068] In an optional implementation manner, the data fusion module 21 is specifically configured to: Perform adaptive bilateral filtering processing on the collected original vision data according to a preset spatial distance standard deviation, pixel value difference standard deviation, and normalization coefficient to obtain denoised vision data; Perform coordinate deviation alignment processing on the denoised vision data through a preset checkerboard calibration model to obtain aligned and optimized vision data; Perform target detection processing on the aligned and optimized vision data through a preset YOLOv7 model to obtain a target object bounding box and semantic labels; Perform depth information mapping processing on the aligned and optimized vision data according to a preset camera internal parameter to obtain spatial point cloud data, and perform data encapsulation processing on the spatial point cloud data, the target object bounding box, and the semantic labels according to a preset data format to obtain the three-dimensional semantic point cloud data.
[0069] The coordinate optimization module 22 is configured to perform coordinate system alignment processing on the three-dimensional semantic point cloud data according to a preset robot base coordinate system to obtain a set of target pose parameters in the robot base coordinate system.
[0070] In an optional implementation manner, the coordinate optimization module 22 is specifically configured to: Perform least squares calculation according to a preset robot base coordinate system through the checkerboard calibration model to obtain a camera-robot coordinate system transformation matrix; Perform transformation processing on the object coordinate data in the three-dimensional semantic point cloud data according to the camera-robot coordinate system transformation matrix to obtain target coordinate data in the robot base coordinate system; Perform principal component analysis processing on the target coordinate data to obtain the object attitude angle data.
[0071] The path planning module 23 is used to perform dynamic path planning on the target pose parameter set to obtain a joint space motion trajectory.
[0072] In an optional embodiment, the path planning module 23 is specifically configured to: Construct an occupancy grid map for the target pose parameter set to obtain an obstacle marked area; Perform a global path search on the target pose parameter set according to a preset improved A* algorithm and the obstacle marked area to obtain optimal discrete path points; Perform B-spline curve fitting on the optimal discrete path points according to a preset acceleration constraint to obtain the joint space motion trajectory.
[0073] The instruction synthesis module 24 is used to perform inverse kinematics solution and motor instruction synthesis on the joint space motion trajectory according to preset robot D-H model parameters to obtain a motor control instruction.
[0074] In an optional embodiment, the instruction synthesis module 24 is specifically configured to: Construct a kinematic model according to preset robot D-H model parameters to obtain a forward kinematic formula; Perform pose mapping on the joint space motion trajectory according to the forward kinematic formula to obtain end-effector pose data; Construct an objective function according to preset target end-effector pose data and the end-effector pose data, and perform inverse kinematics iterative solution on the objective function using the Levenberg–Marquardt algorithm to obtain each joint angle; Convert each joint angle into a motor pulse signal according to a preset encoder pulse number and reduction ratio, and perform synthesis processing on the motor pulse signals according to the time sequence to obtain the motor control instruction.
[0075] The feedback optimization module 25 is used to perform vision-based servo and impedance control closed-loop feedback optimization on the motor control instruction according to the end-effector state data collected in real time to obtain a dynamically corrected control instruction.
[0076] In an optional embodiment, the feedback optimization module 25 is specifically configured to: Perform target detection on the real-time vision data to obtain a target feature point set; Calculate a feature point deviation according to the target end-effector pose data and the target feature point set, and perform a pseudo-inverse calculation of the Jacobian matrix on the feature point deviation to obtain a joint angle correction amount; Adjust the end - position deviation of the contact - force data through a preset impedance - control model to obtain a force - control position correction amount; Superpose and optimize the motor - control instruction according to the joint - angle correction amount and the force - control position correction amount to obtain the dynamic - correction control instruction.
[0077] In an alternative embodiment, the vision - based collaborative - robot end - control device 2 further includes a dynamic - monitoring module 26, and the dynamic - monitoring module 26 is specifically configured to: Perform human - pose tracking detection on the real - time vision data to obtain skeletal key points; Calculate the distance based on the received robot - surface key points and the skeletal key points to obtain the minimum distance; Compare the minimum distance with a preset distance threshold to obtain a corresponding emergency - response signal.
[0078] It should be understood that the various change modes and specific embodiments in the methods provided in the above - mentioned embodiments are equally applicable to the vision - based collaborative - robot end - control device in this embodiment. Through the foregoing detailed description of the vision - based collaborative - robot end - control method, those skilled in the art can clearly know the implementation method of the vision - based collaborative - robot end - control device in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.
[0079] As Figure 3 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0080] In a preferred embodiment of the present invention, the electronic device 3 may include, but is not limited to: a memory 31, at least one processor 32, and at least one communication bus 33.
[0081] Those skilled in the art should understand that Figure 3 the structure of the electronic device 3 shown does not constitute a limitation on the embodiments of the present invention. The electronic device 3 may further include more or fewer other hardware or software than shown, or different component arrangements.
[0082] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre - set or stored instructions, and its hardware includes, but is not limited to, a microprocessor, an application - specific integrated circuit, a programmable gate array, a digital signal processor, and an embedded device, etc.
[0083] It should be noted that the electronic device 3 is only an example, and other existing or future - emerging electronic products that can be adapted to the present application should also be included in the protection scope of the present application and are incorporated herein by reference.
[0084] In some embodiments, a computer program is stored in the memory 31. When the computer program is executed by the at least one processor 32, all or part of the steps in the described vision-based collaborative robot end control method are implemented. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data. Further, the computer-readable storage medium may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, and the like.
[0085] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3, connecting various components of the entire electronic device 3 through various interfaces and lines. By running or executing the programs or modules stored in the memory 31 and calling the data stored in the memory 31, various functions of the electronic device 3 are executed and data is processed. For example, when the at least one processor 32 executes the computer program stored in the memory 31, all or part of the steps in the vision-based collaborative robot end control method described in the embodiments of the present application are implemented; or all or part of the functions of the vision-based collaborative robot end control device are implemented. The at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips.
[0086] In some embodiments, the at least one communication bus 33 is arranged to enable connection communication between the memory 31 and the at least one processor 32, etc. Although not shown, the electronic device 3 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 32 through a power management device, so as to realize functions such as management of charging, discharging, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 3 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0087] The integrated units implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing an electronic device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute a part of the methods described in various embodiments of the present application.
[0088] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0089] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0090] The above are all preferred embodiments of the present application. The protection scope of the present application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A vision-based end control method for collaborative robots, characterized in that, The method includes: Performing multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original visual data to obtain three-dimensional semantic point cloud data; Performing coordinate system alignment processing on the three-dimensional semantic point cloud data according to a preset robot base coordinate system to obtain a set of target pose parameters in the robot base coordinate system; Performing dynamic path planning processing on the set of target pose parameters to obtain a joint space motion trajectory; Performing inverse kinematics solution and motor command synthesis processing on the joint space motion trajectory according to preset robot D-H model parameters to obtain motor control commands; Performing vision-based servo and impedance control closed-loop feedback optimization processing on the motor control commands according to the real-time collected end effector state data to obtain dynamically corrected control commands.
2. The vision-based end control method for a collaborative robot according to claim 1, wherein The performing multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original visual data to obtain three-dimensional semantic point cloud data includes: Performing adaptive bilateral filtering processing on the collected original visual data according to preset spatial distance standard deviation, pixel value difference standard deviation, and normalization coefficient to obtain denoised visual data; Performing coordinate deviation alignment processing on the denoised visual data through a preset checkerboard calibration model to obtain aligned and optimized visual data; Performing target detection processing on the aligned and optimized visual data through a preset YOLOv7 model to obtain target object bounding boxes and semantic labels; Performing depth information mapping processing on the aligned and optimized visual data according to preset camera internal parameters to obtain spatial point cloud data, and performing data encapsulation processing on the spatial point cloud data, the target object bounding boxes, and the semantic labels according to a preset data format to obtain the three-dimensional semantic point cloud data.
3. The vision-based end control method for a collaborative robot according to claim 2, wherein The set of target pose parameters in the robot base coordinate system includes target coordinate data and object attitude angle data. The performing coordinate system alignment processing on the three-dimensional semantic point cloud data according to a preset robot base coordinate system to obtain the set of target pose parameters in the robot base coordinate system includes: Calculating the camera-robot coordinate system transformation matrix through the checkerboard calibration model according to a preset robot base coordinate system by the least squares method; Performing transformation processing on the object coordinate data in the three-dimensional semantic point cloud data according to the camera-robot coordinate system transformation matrix to obtain the target coordinate data in the robot base coordinate system; Performing principal component analysis processing on the target coordinate data to obtain the object attitude angle data.
4. The vision-based end control method for a collaborative robot according to claim 1, wherein The performing dynamic path planning processing on the set of target pose parameters to obtain a joint space motion trajectory includes: Constructing an occupancy grid map for the set of target pose parameters to obtain an obstacle marked area; Performing global path search on the set of target pose parameters according to a preset improved A* algorithm and the obstacle marked area to obtain optimal discrete path points; Performing B-spline curve fitting processing on the optimal discrete path points according to a preset acceleration constraint to obtain the joint space motion trajectory.
5. The vision-based end control method for a collaborative robot according to claim 1, characterized in that Performing inverse kinematics solution and motor command synthesis processing on the joint space motion trajectory according to the preset robot D-H model parameters to obtain motor control commands includes: Constructing a kinematic model according to the preset robot D-H model parameters to obtain a forward kinematics formula; Performing pose mapping on the joint space motion trajectory according to the forward kinematics formula to obtain end-effector pose data; Constructing an objective function based on the preset target end-effector pose data and the end-effector pose data, and performing inverse kinematics iterative solution on the objective function using the Levenberg–Marquardt algorithm to obtain each joint angle; Converting each joint angle into a motor pulse signal according to the preset encoder pulse number and reduction ratio, and performing synthesis processing on the motor pulse signals according to the time sequence to obtain the motor control commands.
6. The vision-based end control method for a collaborative robot according to claim 5, wherein The end-effector state data includes real-time visual data and contact force data. Optimizing the motor control commands based on vision servo and impedance control closed-loop feedback according to the real-time collected end-effector state data to obtain dynamic correction control commands includes: Performing target detection processing on the real-time visual data to obtain a set of target feature points; Calculating the feature point deviation based on the target end-effector pose data and the set of target feature points, and performing pseudo-inverse calculation of the Jacobian matrix on the feature point deviation to obtain a joint angle correction amount; Adjusting the end-effector position deviation of the contact force data through a preset impedance control model to obtain a force control position correction amount; Performing superposition optimization processing on the motor control commands according to the joint angle correction amount and the force control position correction amount to obtain the dynamic correction control commands.
7. The method for controlling the end of a collaborative robot based on vision according to claim 6, wherein The method further includes: Performing human pose tracking detection on the real-time visual data to obtain skeletal key points; Calculating the distance based on the received key points on the robot surface and the skeletal key points to obtain the minimum distance; Comparing the minimum distance with a preset distance threshold to obtain a corresponding emergency response signal.
8. A vision-based end control device for a collaborative robot, characterized in that, The device includes: A data fusion module for performing multi-modal dynamic anti-interference filtering and semantic fusion processing on the collected original visual data to obtain three-dimensional semantic point cloud data; A coordinate optimization module for performing coordinate alignment processing on the three-dimensional semantic point cloud data according to the preset robot base coordinate system to obtain a set of target pose parameters in the robot base coordinate system; A path planning module for performing dynamic path planning processing on the set of target pose parameters to obtain a joint space motion trajectory; An instruction synthesis module for performing inverse kinematics solution and motor instruction synthesis processing on the joint space motion trajectory according to the preset robot D-H model parameters to obtain motor control commands; A feedback optimization module for optimizing the motor control commands based on vision servo and impedance control closed-loop feedback according to the real-time collected end-effector state data to obtain dynamic correction control commands.
9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the vision-based collaborative robot end control method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the vision-based collaborative robot end control method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Modularized movement control system of tool changing robot
CN120697047A
High-precision self-adaptive assembling and adjusting system for automobile door cover based on physical field digital twinning
CN120848438A
Humanoid robot control method based on multi-modal biofeedback and related equipment
CN121132640A
Robot double-source joint information impedance control method and system, medium and robot
CN121132661A
Self-adaptive assembly guiding method and system based on multi-source data fusion
CN121189776A