A dynamic visual guidance method and system for multi-robot collaborative assembly

CN121821375BActive Publication Date: 2026-08-07SUZHOU ZHENGTU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU ZHENGTU INTELLIGENT TECH CO LTD
Filing Date
2026-01-19
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

然而,在实际汽车仪表板装配等场景中,车身薄壁件的柔性变形与仪表表面的动态反光,使得基于预标定参数的视觉引导系统面临挑战,工件形变导致实际几何特征偏离理论模型,复杂光照条件下的反光干扰影响视觉智能算法的特征提取稳定性,因此现有技术存在不足

Benefits of technology

[0047] Based on the observation dataset and the initial hand-eye matrix, this invention obtains the correction amount of the hand-eye matrix through the forward projection principle and nonlinear optimization. According to the reprojection error, unsupervised clustering and spatial interpolation are used to generate the updated hand-eye matrix and local pixel compensation field, respectively. This realizes dynamic adaptive compensation for the flexible deformation of the vehicle body and the dynamic reflection of the instrument, solves the problem of dynamic calibration failure caused by changes in working conditions, and improves the accuracy of multi-robotic arm collaborative assembly and the system's adaptability in complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121821375B_ABST
    Figure CN121821375B_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic visual guidance method and system of multi-robot arm collaborative assembly, belong to industrial robot visual guidance technical field, its technical scheme main points include: through binocular vision system synchronous acquisition image sequence, through robot controller synchronous record the pose sequence of each mechanical arm, obtain observation dataset;Through optical flow tracking algorithm obtains the two-dimensional trajectory flow of bodywork assembly area feature point and instrument body feature point, and obtains the hand-eye matrix correction amount by nonlinear optimization algorithm;Based on the re-projection error of each feature point, the updated hand-eye matrix and local pixel compensation field are obtained by space interpolation algorithm;Through three-dimensional reconstruction and coordinate compensation, the compensated assembly target coordinate is obtained, the application is according to re-projection error and is used to realize the adaptive compensation of bodywork flexible deformation and instrument dynamic reflection by unsupervised clustering and space interpolation, improve the precision of multi-robot arm collaborative assembly and the adaptive ability of system in complex industrial environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vision guidance technology for industrial robots, and more specifically to a dynamic vision guidance method and system for collaborative assembly of multiple robotic arms. Background Technology

[0002] In the field of industrial robot technology, vision-based guidance systems have become a core technology for achieving collaborative assembly of multiple robotic arms. Industrial vision systems establish coordinate transformation relationships between the camera and the robotic arm through hand-eye calibration, providing the robot with precise target localization. Traditional methods typically use these transformation parameters after calibration in a controlled environment, a static calibration model based on the ideal assumptions of invariant workpiece geometry and a stable imaging environment. However, in real-world scenarios such as automotive dashboard assembly, the flexible deformation of thin-walled body parts and the dynamic reflections on instrument surfaces pose challenges to vision guidance systems based on pre-calibrated parameters. Workpiece deformation causes actual geometric features to deviate from the theoretical model, and reflection interference under complex lighting conditions affects the feature extraction stability of visual intelligence algorithms. Therefore, existing technologies have shortcomings. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to provide a dynamic visual guidance method and system for multi-robotic arm collaborative assembly. By constructing an observation dataset and extracting the two-dimensional trajectory flow of feature points, the hand-eye matrix correction is obtained based on the forward projection principle and nonlinear optimization. The updated hand-eye matrix and local pixel compensation field are generated according to unsupervised clustering and spatial interpolation, respectively. This achieves dynamic adaptive compensation for vehicle body flexible deformation and instrument dynamic reflection, solves the problem of dynamic calibration failure caused by changes in working conditions, and improves the accuracy of multi-robotic arm collaborative assembly and the system's adaptability in complex industrial environments.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] This invention provides a dynamic vision-guided method for collaborative assembly of multiple robotic arms, comprising:

[0006] Based on the initial hand-eye matrix control, the assembly robot arm and the tightening robot arm move collaboratively from their respective initial estimated poses to the corresponding reference target points. The image sequence is synchronously acquired through the binocular vision system, and the pose sequence of each robot arm is synchronously recorded through the robot controller to obtain the observation dataset.

[0007] Based on the observation dataset, the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained by optical flow tracking algorithm. The predicted motion trajectory of each feature point is obtained by forward projection principle, and the hand-eye matrix correction is obtained by nonlinear optimization algorithm.

[0008] The reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow. Based on the reprojection error, the updated hand-eye matrix and local pixel compensation field are obtained through unsupervised clustering algorithm and spatial interpolation algorithm.

[0009] Based on the updated hand-eye matrix and local pixel compensation field, the compensated assembly target coordinates are obtained through three-dimensional reconstruction and coordinate compensation. The assembly target coordinates include the vehicle body assembly coordinates and the instrument installation coordinates.

[0010] As a further improvement of the present invention, based on the observation dataset, a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained through an optical flow tracing algorithm; the predicted motion trajectory of each feature point is obtained through the forward projection principle; and the hand-eye matrix correction is obtained through a nonlinear optimization algorithm, including:

[0011] Based on the image sequence of the observation dataset, a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained through an optical flow tracing algorithm;

[0012] Based on the pose sequence, two-dimensional trajectory flow, and initial hand-eye matrix of the observation dataset, the predicted motion trajectory of each feature point is obtained through the forward projection principle, and the correction amount of the hand-eye matrix is ​​obtained through a nonlinear optimization algorithm.

[0013] As a further improvement of the present invention, the step of obtaining a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument panel body based on the image sequence of the observation dataset using an optical flow tracing algorithm includes:

[0014] Based on the image sequence of the observation dataset, the initial feature points of the vehicle body assembly area and the instrument body area are extracted;

[0015] Based on the initial feature points, the motion trajectory of the feature points is obtained by tracking the feature points in consecutive image frames using an optical flow tracing algorithm.

[0016] Based on the feature point matching relationship between the left and right eye images, a trajectory filtering algorithm is used to remove mismatched points, resulting in a stable two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body.

[0017] As a further improvement of the present invention, the step of obtaining the predicted motion trajectory of each feature point based on the pose sequence, two-dimensional trajectory flow, and initial hand-eye matrix of the observation dataset through the forward projection principle, and obtaining the hand-eye matrix correction amount through a nonlinear optimization algorithm, includes:

[0018] For each feature point in the two-dimensional trajectory flow, based on the pose sequence of the observation dataset and the initial hand-eye matrix, the predicted motion trajectory of each feature point is obtained through the forward projection principle.

[0019] Based on the predicted motion trajectory of each feature point, the hand-eye matrix correction is obtained through a nonlinear optimization algorithm.

[0020] As a further improvement of the present invention, the reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow, and the updated hand-eye matrix and local pixel compensation field are obtained based on the reprojection error through an unsupervised clustering algorithm and a spatial interpolation algorithm, including:

[0021] The reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow.

[0022] Based on the reprojection error, an unsupervised clustering algorithm is used to obtain a set of rigid feature points and a set of reflective feature points;

[0023] The updated hand-eye matrix is ​​obtained based on the rigid feature point set and its reprojection error.

[0024] Based on the set of reflective feature points and their reprojection errors, a local pixel compensation field is constructed using a spatial interpolation algorithm.

[0025] As a further improvement of the present invention, the reprojection error of each feature point obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow includes:

[0026] Based on the aforementioned hand-eye matrix correction, the corrected hand-eye matrix is ​​obtained through Lie algebra integration.

[0027] Based on the corrected hand-eye matrix and the pose sequence of the observation dataset, the recalculated predicted position of each feature point is obtained through the forward projection principle.

[0028] Based on the recalculated predicted position and the actual observed position in the two-dimensional trajectory flow, the reprojection error of each feature point is calculated using Euclidean distance.

[0029] As a further improvement of the present invention, the step of obtaining the rigid feature point set and the reflective feature point set based on the reprojection error using an unsupervised clustering algorithm includes:

[0030] Based on the reprojection error and spatial coordinates of all feature points, preliminary clustering results are obtained by performing preliminary clustering using an unsupervised clustering algorithm;

[0031] Based on the preliminary clustering results, the optimized clustering boundaries are obtained using a region growing algorithm.

[0032] Based on the optimized clustering results, the feature points are divided into a set of rigid feature points with continuous spatial distribution and a set of reflective feature points with discrete spatial distribution using statistical analysis methods.

[0033] As a further improvement of the present invention, the step of constructing a local pixel compensation field based on the set of reflective feature points and their reprojection errors using a spatial interpolation algorithm includes:

[0034] Based on the spatial coordinates and reprojection error of the reflective feature point set, the pixel offset of each feature point is calculated by weighted averaging.

[0035] Based on the pixel offset and spatial distribution of the reflective feature points, an initial compensation field is established using a radial basis function interpolation algorithm.

[0036] Based on the requirement of continuity of the compensation field, the smoothness of the compensation field is optimized by thin plate spline spatial interpolation algorithm to obtain the final local pixel compensation field.

[0037] As a further improvement of the present invention, the compensated assembly target coordinates are obtained through three-dimensional reconstruction and coordinate compensation based on the updated hand-eye matrix and local pixel compensation field. The assembly target coordinates include vehicle body assembly coordinates and instrument installation coordinates, including:

[0038] Based on real-time acquired binocular images, the initial pixel coordinates of vehicle body feature points and instrument panel feature points are obtained through feature extraction algorithms;

[0039] Based on the initial pixel coordinates of the instrument feature points, a local pixel compensation field is applied, and the compensated pixel coordinates of the instrument feature points are obtained through coordinate mapping.

[0040] Based on the compensated pixel coordinates of the instrument feature points, the instrument installation coordinates are calculated using the binocular triangulation principle.

[0041] Based on the initial pixel coordinates of the vehicle body feature points and the updated hand-eye matrix, the vehicle body assembly coordinates are calculated using the binocular triangulation principle.

[0042] This invention provides a dynamic vision guidance system for multi-robotic arm collaborative assembly, the dynamic vision guidance system for multi-robotic arm collaborative assembly comprising:

[0043] Data acquisition module: Based on the initial hand-eye matrix, the assembly robot arm and the tightening robot arm are controlled to move collaboratively from their respective initial estimated poses to the corresponding reference target points. The image sequence is acquired synchronously through the binocular vision system, and the pose sequence of each robot arm is recorded synchronously through the robot controller to obtain the observation dataset.

[0044] Visual analysis module: Based on the observation dataset, the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained through optical flow tracking algorithm, the predicted motion trajectory of each feature point is obtained through forward projection principle, and the hand-eye matrix correction amount is obtained through nonlinear optimization algorithm;

[0045] Error compensation module: Based on the hand-eye matrix correction and the two-dimensional trajectory flow, the reprojection error of each feature point is obtained. Based on the reprojection error, the updated hand-eye matrix and local pixel compensation field are obtained through unsupervised clustering algorithm and spatial interpolation algorithm.

[0046] Coordinate generation module: Based on the updated hand-eye matrix and local pixel compensation field, the compensated assembly target coordinates are obtained through three-dimensional reconstruction and coordinate compensation. The assembly target coordinates include vehicle body assembly coordinates and instrument installation coordinates.

[0047] Based on the observation dataset and the initial hand-eye matrix, this invention obtains the correction amount of the hand-eye matrix through the forward projection principle and nonlinear optimization. According to the reprojection error, unsupervised clustering and spatial interpolation are used to generate the updated hand-eye matrix and local pixel compensation field, respectively. This realizes dynamic adaptive compensation for the flexible deformation of the vehicle body and the dynamic reflection of the instrument, solves the problem of dynamic calibration failure caused by changes in working conditions, and improves the accuracy of multi-robotic arm collaborative assembly and the system's adaptability in complex industrial environments. Attached Figure Description

[0048] Figure 1 This is a flowchart of the steps of the dynamic vision guidance method for multi-robotic arm collaborative assembly of the present invention.

[0049] Figure 2 This is a schematic diagram illustrating the steps of obtaining the updated hand-eye matrix and local pixel compensation field based on the reprojection error.

[0050] Figure 3 This is a schematic diagram illustrating the steps involved in obtaining the compensated coordinates of the assembly target through 3D reconstruction and coordinate compensation.

[0051] Figure 4 This is a schematic diagram of the dynamic vision guidance system for multi-robotic arm collaborative assembly according to the present invention. Detailed Implementation

[0052] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof.

[0053] Identical parts are indicated by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "up," and "down" used in the following description refer to directions in the accompanying drawings, while the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific part, respectively.

[0054] The term "and / or" in the following text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0055] like Figure 1 As shown, the present invention provides a dynamic vision guidance method for collaborative assembly of multiple robotic arms, comprising:

[0056] Based on the initial hand-eye matrix control, the assembly robot arm and the tightening robot arm move collaboratively from their respective initial estimated poses to the corresponding reference target points. The image sequence is synchronously acquired through the binocular vision system, and the pose sequence of each robot arm is synchronously recorded through the robot controller to obtain the observation dataset.

[0057] Based on the observation dataset, the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained by optical flow tracking algorithm. The predicted motion trajectory of each feature point is obtained by forward projection principle, and the hand-eye matrix correction is obtained by nonlinear optimization algorithm.

[0058] The reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow. Based on the reprojection error, the updated hand-eye matrix and local pixel compensation field are obtained through unsupervised clustering algorithm and spatial interpolation algorithm.

[0059] Based on the updated hand-eye matrix and local pixel compensation field, the compensated assembly target coordinates are obtained through 3D reconstruction and coordinate compensation. The assembly target coordinates include the vehicle body assembly coordinates and the instrument installation coordinates.

[0060] The initial hand-eye matrix represents the transformation relationship between the camera coordinate system and the robot tool coordinate system, including the transformation matrix from the camera to the end effector of the robotic arm and the transformation relationship between the dual robotic arm base coordinate systems. It is obtained by controlling the robotic arm to move the camera from a preset pose to observe the fixed calibration plate, acquiring images and robot joint angles at each pose, and solving the problem using a hand-eye calibration algorithm. The assembly robotic arm is a six-DOF industrial robot that performs instrument grasping and placement operations; the tightening robotic arm is a six-DOF industrial robot that performs bolt tightening operations. The initial estimated pose is the starting pose of the robotic arm's initial movement, derived based on the ideal assembly position of a standard, deformation-free vehicle body and a non-reflective instrument panel. This is achieved through calibration. The system controls the robotic arm to slowly approach the assembly position. The position is derived by precisely locating feature points using a vision system and recording the corresponding Cartesian coordinates of the robot. The reference target point is the ideal three-dimensional coordinates of the body assembly hole and the instrument mounting point. In offline mode, using a standard, non-deformable body, non-reflective instruments, and a calibrated system, the robotic arm is controlled to slowly approach the assembly position. The optimal assembly position is obtained by precisely locating feature points using a vision system and recording the corresponding Cartesian coordinates of the robot. The binocular vision system is a vision acquisition device containing left and right cameras. It is used to synchronously acquire image sequences of the assembly area at a fixed preset frequency and achieve hardware synchronization with the robot controller to ensure data time alignment. It is used to acquire workpiece images. The image sequence is time-series... The system consists of a set of multiple frames of binocular images arranged in a sequential order, including images captured by the left and right cameras at the same time. The robot controller is an industrial controller with real-time motion control capabilities, used to record the robot arm's pose. The robot arm's pose sequence is a set of pose transformation matrices of the robot arm's end effector in a base coordinate system, recorded in chronological order. The observation dataset is a time-synchronized causal observation window dataset, including the aforementioned image and pose sequences. It is obtained by controlling the assembly and tightening robots to perform linear approximation motion at a preset speed. The image sequences are acquired by the binocular vision system, the pose sequences are recorded by the robot controller, and then processed synchronously by hardware. This dataset is used to establish the correspondence between the robot arm's motion and image changes. The optical flow tracking algorithm is a dense optical flow algorithm. The algorithm is used to extract the motion trajectories of feature points in an image sequence; the two-dimensional trajectory flow of feature points in the body assembly area is the set of motion trajectories of feature points around the body assembly holes in the image coordinate system; the two-dimensional trajectory flow of feature points in the instrument body is the set of motion trajectories of feature points on the instrument surface in the image coordinate system; the forward projection principle is a mathematical method for projecting three-dimensional points onto a two-dimensional image plane based on the camera imaging model, and projects three-dimensional spatial points onto a two-dimensional image plane based on the robot arm pose and the initial hand-eye matrix, which is used to calculate the predicted motion trajectory of each feature point; the predicted motion trajectory is the theoretical motion path of the feature point in the image calculated based on robot kinematics and the hand-eye matrix, reflecting the pixel position change of the feature point as the robot arm moves under ideal conditions;The nonlinear optimization algorithm is used to construct the reprojection error objective function and iteratively optimize the solution for the hand-eye matrix correction. The optimal solution for the hand-eye matrix correction is obtained by minimizing the reprojection error, where the correction is a Lie algebraic form of the correction. The reprojection error is the Euclidean distance between the predicted image coordinates and the actual observed image coordinates of the feature points. The unsupervised clustering algorithm is a density clustering algorithm that classifies feature points according to their spatial distribution patterns based on the reprojection error residuals. It identifies spatially continuous rigid feature point sets and spatially discrete reflective feature point sets, achieving separation of geometric deformation errors and optical interference errors. The spatial interpolation algorithm is a thin-plate spline interpolation algorithm. A continuous local pixel compensation field is constructed based on the pixel offset of reflective feature points; the updated hand-eye matrix is ​​the camera and robotic arm tool coordinate system transformation matrix after correction and optimization of the hand-eye matrix; the local pixel compensation field is a two-dimensional pixel offset mapping function constructed based on the offset of reflective feature points, used to compensate for local optical interference errors; 3D reconstruction and coordinate compensation is a process of obtaining 3D coordinates through binocular triangulation and coordinate transformation and performing error compensation, used to obtain accurate assembly positions; the assembly target coordinates are the final assembly position coordinates after dynamic compensation, including vehicle body assembly coordinates and instrument installation coordinates, used for multi-robotic arm collaborative assembly, providing accurate position references for the collaborative movement of robotic arms.

[0061] This embodiment establishes the correspondence between motion and image by dynamically acquiring time-synchronized observation datasets, corrects the hand-eye matrix through feature trajectory flow extraction and nonlinear optimization, separates error sources through unsupervised clustering and constructs a compensation field by combining spatial interpolation, and obtains accurate assembly target coordinates through 3D reconstruction and coordinate compensation. This achieves dynamic adaptive compensation for vehicle body flexible deformation and instrument reflection interference, improves the sub-millimeter assembly accuracy of multi-robotic arm collaborative assembly, and enhances the stability, accuracy, adaptability, and system's adaptive capability and robustness in the assembly process.

[0062] Furthermore, this embodiment provides a step for determining the initial hand-eye matrix and reference target points, including: performing a system initialization phase to establish a coordinate transformation reference and an ideal assembly position reference between the vision system and the robot system. The system initialization phase includes two key processes: hand-eye system calibration for the multi-arm robotic system and establishment of ideal assembly reference points. The hand-eye system calibration step is used to obtain the initial hand-eye matrix between the binocular vision system and each robotic arm, the relative transformation relationship between the base coordinate systems of each robotic arm, and the camera intrinsic parameters and distortion coefficients. The ideal assembly reference point establishment step is used to obtain the ideal three-dimensional coordinates of the vehicle body mounting holes and instrument mounting points and construct an ideal assembly point database. Specifically, the multi-arm hand-eye system calibration process is based on a calibration board, a binocular camera, an assembly robot arm, and a tightening robot arm. By controlling each robot arm to drive the camera to observe the fixed calibration board from a preset pose, binocular images and robot arm joint angles are acquired at each pose. The hand-eye calibration algorithm is used to solve the hand-eye equation to obtain the initial hand-eye matrix, as well as the camera intrinsic parameters K and distortion coefficients D. The initial hand-eye matrix includes the transformation matrix and the transformation relationship of the dual robot arm base coordinate system. The ideal assembly reference point establishment process is based on a standard non-deformable car body, non-reflective instruments, and the execution of a system that has completed hand-eye calibration. By controlling the robot arm to slowly approach the preset assembly position, the vision system is used to accurately locate the feature points of the car body assembly area and the instrument body feature points. The robot's Cartesian coordinates corresponding to the optimal assembly position are recorded, and the ideal coordinates of the car body assembly holes and the instrument mounting points are obtained and stored as an ideal assembly point database.

[0063] This embodiment establishes multi-coordinate system transformation relationships through hand-eye calibration and determines the ideal assembly datum through visual guidance teaching, providing accurate initial parameters and assembly quality evaluation standards for subsequent dynamic calibration.

[0064] Furthermore, this embodiment provides a method for obtaining a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument panel based on an observation dataset using an optical flow tracing algorithm, obtaining the predicted motion trajectory of each feature point through forward projection, and obtaining the hand-eye matrix correction amount through a nonlinear optimization algorithm, including:

[0065] Based on the image sequence of the observation dataset, the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained by optical flow tracing algorithm;

[0066] Based on the pose sequence, two-dimensional trajectory flow and initial hand-eye matrix of the observation dataset, the predicted motion trajectory of each feature point is obtained through the forward projection principle, and the correction amount of the hand-eye matrix is ​​obtained through a nonlinear optimization algorithm.

[0067] Specifically, based on the image sequence of the observation dataset, Harris corner points are detected in the first frame image as initial feature points. Harris corner points are a set of pixels with significant gray-level changes and obvious edge features. A total of [number missing] Harris corner points were detected. Initial feature points are identified, and then the dense optical flow algorithm is used to track these initial feature points in consecutive frames of the image sequence. The reliability of each feature point's trajectory is verified by left and right eye image matching using a binocular vision system, and feature points with trajectory jumps greater than a preset pixel threshold are removed. Anomalies are identified, and a two-dimensional trajectory flow of feature points in the vehicle assembly area is obtained. Two-dimensional trajectory flow of instrument body feature points Then, based on the pose sequence, 2D trajectory flow, and initial hand-eye matrix of the observation dataset, for each feature point, its predicted image position is calculated using the forward projection principle. The forward projection principle is a mathematical method that projects 3D spatial points onto a 2D image plane according to the camera imaging model. The formula for calculating the predicted image position is as follows: ,in For projection function, for The pose transformation matrix of the robotic arm's end effector relative to the base coordinate system at any given time. Let be the transformation matrix from the camera to the robotic arm end effector in the initial hand-eye matrix. Let be the 3D spatial coordinates of the feature points, and let the predicted motion trajectory be the set of predicted image positions of each feature point in consecutive frames. Then, a reprojection error objective function is constructed with the goal of minimizing the reprojection error, where the reprojection error is the Euclidean distance between the predicted image position and the observed image position of the feature point. The reprojection error objective function is: ,in This is the correction value for the hand-eye matrix. The observed image positions of the feature points are used. Finally, the LM algorithm is used to iteratively optimize the objective function of the reprojection error. By adjusting the hand-eye matrix correction, the sum of the reprojection errors is minimized, and the optimal hand-eye matrix correction is obtained. The hand-eye matrix correction is a 6-dimensional Lie algebra vector used to quantify the calibration error of the initial hand-eye matrix.

[0068] For example, this embodiment assumes that the observation dataset contains The image sequence and pose sequence at each time sampling point are obtained through optical flow tracing. Individual vehicle body feature points and Two-dimensional trajectory flow of instrument feature points; based on the robotic arm pose sequence and initial hand-eye matrix, the predicted motion trajectory of each feature point in the image coordinate system is calculated using the forward projection principle; a reprojection error objective function is constructed and the hand-eye matrix correction in 6-dimensional Lie algebra form is obtained through nonlinear optimization. .

[0069] This embodiment extracts accurate dynamic feature trajectory flow through corner detection and dense optical flow tracking, constructs feature point prediction motion trajectory through the forward projection principle, establishes a visual and motion consistency optimization model through the reprojection error objective function, and solves the optimal hand-eye matrix correction amount through the LM algorithm iterative solution. This achieves accurate identification and quantification of the initial calibration error, improves the reliability of the feature trajectory and the adaptability of the hand-eye matrix, and provides high-precision basic data support for subsequent error separation and compensation.

[0070] Furthermore, this embodiment provides a step for obtaining the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument panel body based on an image sequence of an observation dataset using an optical flow tracing algorithm, including:

[0071] Based on the image sequence of the observation dataset, the initial feature points of the vehicle body assembly area and the instrument body area are extracted;

[0072] Based on the initial feature points, the motion trajectory of the feature points is obtained by tracking the feature points in consecutive image frames using an optical flow tracing algorithm.

[0073] Based on the feature point matching relationship of the vehicle body in the left and right eye images, a trajectory filtering algorithm is used to remove mismatched points, resulting in a stable two-dimensional trajectory flow of the assembly area feature points and the instrument body feature points.

[0074] The initial feature points are a set of corner points with significant gray-level changes in the image, extracted from the first frame of the observation dataset using the Harris corner detection algorithm. The optical flow tracing algorithm is a dense optical flow algorithm that calculates the motion vectors of all pixels in the image based on the principle of continuous pixel gray-level changes. It is used to accurately track the positional changes of the initial feature points in consecutive image frames, i.e., to trace the motion trajectory of the feature points in consecutive image frames. The motion trajectory of the feature points is the set of paths through which the pixel coordinates of the initial feature points change over time in each consecutive frame of the image sequence in the observation dataset. It is obtained by calculating the coordinates of the initial feature points in each frame using the optical flow tracing algorithm. This is used to reflect the dynamic motion patterns of feature points; the feature point matching relationship between the left and right eye images is the correspondence between the same spatial feature points in the left and right camera images acquired by the binocular vision system, obtained by calculating the gray-level histogram similarity and spatial position constraints of the feature points, and is used to verify the spatial consistency of the feature point trajectory; the trajectory filtering algorithm is an outlier removal algorithm based on trajectory stability judgment, used to remove mismatched points in the motion trajectory that do not conform to spatial consistency and temporal continuity; the two-dimensional trajectory flow is a set of stable feature point motion trajectories obtained after trajectory filtering, including the two-dimensional trajectory flow of feature points in the vehicle assembly area and the two-dimensional trajectory flow of feature points in the instrument body.

[0075] Specifically, based on the image sequence of the observation dataset, the left and right images of the first frame of the sequence are selected as reference frames. The Harris corner detection algorithm is used to extract initial feature points in the vehicle assembly area and the instrument panel area of ​​the reference frames, respectively, and a total of [number missing] feature points are detected. There are 10 initial feature points, of which the initial feature points in the vehicle assembly area are 10. The initial feature points of the instrument body area are: individual and Each initial feature point records its pixel coordinates in the left and right eye images; then, based on the extracted... Using an initial feature point, a dense optical flow algorithm is applied to track feature points across consecutive frames of the observed dataset image sequence. By calculating the grayscale gradient and displacement vector between adjacent frames, the pixel coordinates of each initial feature point in each frame are obtained, forming... The motion trajectories of preliminary feature points are generated, each containing a sequence of pixel coordinates for that feature point across all consecutive frames. Then, based on the feature point matching relationship between the left and right eye images, the consistency of each preliminary feature point's motion trajectory is verified. The spatial positional deviation between the feature point trajectory in the left eye image and the corresponding feature point trajectory in the right eye image is calculated, and the temporal continuity of the trajectory is analyzed. If the pixel jump between adjacent frames of the trajectory exceeds a preset threshold... Or the spatial position deviation of the corresponding trajectories of the left and right eyes is greater than a preset threshold. If the trajectory is incorrect, it is determined to be an abnormal trajectory corresponding to a mismatched point. Finally, a trajectory filtering algorithm is used to remove all abnormal trajectories, retaining the feature point trajectories that meet the requirements of spatial consistency and temporal continuity, resulting in... A stable trajectory, wherein the two-dimensional trajectory flow of feature points in the vehicle assembly area includes The trajectory, the two-dimensional trajectory flow of the instrument body feature points includes Trajectory, and Less than or equal to , Less than or equal to , Less than or equal to .

[0076] This embodiment accurately extracts initial feature points through Harris corner detection, achieves stable tracking of feature points in consecutive frames through the Farneback dense optical flow algorithm, verifies trajectory consistency through the matching relationship between left and right eye feature points, and eliminates mismatched abnormal points through a trajectory filtering algorithm. This achieves accurate extraction of the two-dimensional trajectory flow of feature points in the vehicle assembly area and the instrument body, improving the spatial consistency, temporal continuity, and reliability of the feature trajectory. It provides high-precision dynamic feature data support for subsequent forward projection prediction and hand-eye matrix correction calculation.

[0077] Furthermore, this embodiment provides a step-by-step approach based on the pose sequence, two-dimensional trajectory flow, and initial hand-eye matrix of the observation dataset. This approach yields the predicted motion trajectory of each feature point through forward projection and obtains the hand-eye matrix correction value through a nonlinear optimization algorithm. The steps include:

[0078] For each feature point in the two-dimensional trajectory flow, the predicted motion trajectory of each feature point is obtained by using the forward projection principle based on the pose sequence of the observation dataset and the initial hand-eye matrix.

[0079] Based on the predicted motion trajectory of each feature point, the correction amount of the hand-eye matrix is ​​obtained through a nonlinear optimization algorithm.

[0080] The predicted motion trajectory is the theoretical motion path of the feature points in the image coordinate system, which is calculated using the forward projection principle. The nonlinear optimization algorithm is the LM algorithm, which is an iterative optimization algorithm that combines gradient descent and Gaussian and Newton methods. It is used to minimize the reprojection error objective function to solve for the optimal hand-eye matrix correction. The hand-eye matrix correction is a transformation matrix correction in the form of a 6-dimensional Lie algebra, which is solved by the nonlinear optimization algorithm and is used to quantify and correct the calibration error of the initial hand-eye matrix.

[0081] Specifically, for each feature point in the two-dimensional trajectory flow, based on the pose sequence of the observation dataset and the camera intrinsic parameters and distortion coefficients in the initial hand-eye matrix, the three-dimensional spatial coordinates P of each feature point are calculated using a binocular triangulation algorithm. The binocular triangulation algorithm is an algorithm that recovers the three-dimensional spatial coordinates using the disparity information of the left and right eye images of a binocular vision system. A total of [number missing] coordinates are obtained. The three-dimensional spatial coordinates of the feature points, where This represents the total number of feature points in the two-dimensional trajectory flow, including feature points in the vehicle assembly area. Individual and instrument body feature points individual and Then, for each feature point's three-dimensional spatial coordinates P, combined with the pose sequence of the observation dataset at each time step... Given the pose transformation matrix and the initial hand-eye matrix, the feature point is calculated using the forward projection principle. Predicted image location at time Then, the predicted motion trajectory of each feature point is obtained by traversing all time points. The predicted motion trajectory is the sequence of predicted image positions of each feature point in consecutive frames. Then, based on the predicted motion trajectory of each feature point and the sequence of observed image positions in the two-dimensional trajectory stream, a reprojection error objective function is constructed. The reprojection error is the Euclidean distance between the predicted image position and the observed image position of the feature point at each time point. Finally, the initial value of the hand-eye matrix correction is initialized, and the iteration convergence threshold and the maximum number of iterations are set. The reprojection error objective function is iteratively optimized by the LM algorithm. In each iteration, the predicted image position is updated by adjusting the hand-eye matrix correction and the sum of reprojection errors is calculated. The iteration stops when the change in the sum of errors between two consecutive iterations is less than the iteration convergence threshold or the number of iterations reaches the maximum number of iterations, and the optimal hand-eye matrix correction is obtained.

[0082] This embodiment accurately obtains the three-dimensional spatial coordinates of feature points through a binocular triangulation algorithm, constructs the predicted motion trajectory of feature points through the forward projection principle, establishes vision-motion consistency constraints through the reprojection error objective function, and iteratively solves the optimal hand-eye matrix correction amount through the LM algorithm. This achieves accurate quantification and correction of the initial hand-eye matrix calibration error, improves the coordinate transformation accuracy of the vision-guided robot system, provides high-precision parameter support for subsequent error separation and compensation, and ensures the guidance accuracy of multi-arm collaborative assembly.

[0083] Furthermore, this embodiment provides a step of obtaining the reprojection error of each feature point based on the hand-eye matrix correction and the two-dimensional trajectory flow, and obtaining the updated hand-eye matrix and local pixel compensation field based on the reprojection error using an unsupervised clustering algorithm and a spatial interpolation algorithm, including:

[0084] The reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow.

[0085] Based on the reprojection error, an unsupervised clustering algorithm is used to obtain the rigid feature point set and the reflective feature point set;

[0086] The updated hand-eye matrix is ​​obtained based on the rigid feature point set and its reprojection error;

[0087] Based on the set of reflective feature points and their reprojection errors, a local pixel compensation field is constructed using a spatial interpolation algorithm.

[0088] The reprojection error is the Euclidean distance between the predicted image coordinates and the actual observed image coordinates of a feature point. It is obtained by recalculating the predicted position based on the hand-eye matrix correction and comparing it with the actual observed position in the two-dimensional trajectory flow, used to quantify the deviation of the feature point position. The unsupervised clustering algorithm is a density clustering algorithm, which includes the process of classifying feature points based on the spatial distribution pattern of the reprojection error residual, used to identify spatially continuous rigid feature point sets and spatially discrete reflective feature point sets. The rigid feature point set is a set of spatially continuous feature points with small reprojection errors, corresponding to geometric errors such as vehicle body flexible deformation, obtained through the unsupervised clustering algorithm, used for global hand-eye matrix updates. The reflective feature point set is a set of spatially discrete feature points with small reprojection errors. Furthermore, the set of feature points with large reprojection errors corresponds to optical interference errors such as instrument reflections. This set is obtained through an unsupervised clustering algorithm and used for local pixel compensation. The updated hand-eye matrix is ​​the camera and robotic arm tool coordinate system transformation matrix optimized by the rigid feature point set. It is obtained through Lie algebra exponentiation operation of the initial hand-eye matrix and the hand-eye matrix correction, and is used to improve the accuracy of global coordinate transformation. The spatial interpolation algorithm is a thin plate spline interpolation algorithm, which is an interpolation method that maintains data point constraints and minimizes bending energy. It is used to construct a continuous compensation field based on the offset of discrete reflective feature points. The local pixel compensation field is a two-dimensional pixel offset mapping function constructed based on the offset of reflective feature points. It is obtained through a spatial interpolation algorithm and is used to compensate for local optical interference errors.

[0089] Specifically, such as Figure 2 As shown, firstly, based on the correction amount of the hand-eye matrix, the corrected hand-eye matrix is ​​obtained through Lie algebra integration. Then, combining the pose sequence and 2D trajectory flow of the observation dataset, the predicted image position of each feature point is recalculated using the forward projection principle. Next, based on the recalculated predicted position and the actual observed position in the 2D trajectory flow, the reprojection error of each feature point is calculated using Euclidean distance. Then, based on the reprojection errors of all feature points and their spatial coordinates in the image coordinate system, preliminary clustering is performed using an unsupervised clustering algorithm to obtain preliminary clustering results. Finally, based on the preliminary clustering results, the cluster boundaries are optimized using a region growing algorithm to obtain optimized clustering results. Finally, based on the optimized clustering results... The feature points are divided into a set of rigid feature points with continuous spatial distribution and a set of reflective feature points with discrete spatial distribution using statistical analysis methods. Then, based on the rigid feature point set and its reprojection error, the corrected hand-eye matrix is ​​further optimized using a nonlinear optimization algorithm to obtain an updated hand-eye matrix. Finally, based on the spatial coordinates and reprojection error of the reflective feature point set, the pixel offset of each feature point is calculated by weighted averaging. Based on the pixel offsets and spatial distribution of all reflective feature points, an initial compensation field is established using a radial basis function interpolation algorithm. Then, based on the continuity requirement of the compensation field, the smoothness of the compensation field is optimized using a thin plate spline spatial interpolation algorithm to obtain the final local pixel compensation field.

[0090] For example, this embodiment assumes that the two-dimensional trajectory flow contains For each feature point, the error value is calculated through reprojection error. An unsupervised clustering algorithm is then used to set the neighborhood radius to [value missing]. pixels, minimum number of samples Perform clustering to obtain the included A rigid feature point set containing 1 feature point and The set of reflective feature points of feature points, among which The updated hand-eye matrix is ​​obtained by updating the rigid feature point set through Lie algebra. A local pixel compensation field is constructed by thin plate spline interpolation based on the reflective feature point set, covering all pixels in the image area.

[0091] This embodiment accurately calculates reprojection errors based on Lie algebra integration and Euclidean distance. It intelligently separates geometric deformation and optical interference error sources through unsupervised clustering algorithm, optimizes the global hand-eye matrix using rigid feature point set, and constructs a continuous local compensation field through thin plate spline interpolation. This achieves hierarchical compensation for systematic calibration errors and local optical interference errors, improves the accuracy of the hand-eye matrix and the accuracy of local compensation, and enhances the coordinate transformation accuracy and anti-interference capability of multi-robotic arm collaborative assembly.

[0092] Furthermore, this embodiment provides a step for obtaining the reprojection error of each feature point based on the hand-eye matrix correction and the two-dimensional trajectory flow, including:

[0093] Based on the correction amount of the hand-eye matrix, the corrected hand-eye matrix is ​​obtained through Lie algebra integration.

[0094] Based on the corrected hand-eye matrix and the pose sequence of the observation dataset, the recalculated predicted position of each feature point is obtained through the forward projection principle.

[0095] Based on the recalculated predicted position and the actual observed position in the two-dimensional trajectory flow, the reprojection error of each feature point is calculated using Euclidean distance.

[0096] Among them, the Lie algebra integration operation is an exponential mapping operation that converts the Lie algebra vector into a transformation matrix, used to convert the hand-eye matrix correction into a homogeneous transformation matrix, ensuring the orthogonality and accuracy of the correction matrix; the corrected hand-eye matrix is ​​an optimized transformation matrix obtained by applying the hand-eye matrix correction to the initial hand-eye matrix through the Lie algebra integration operation, used to more accurately describe the transformation relationship between the camera and the robotic arm tool coordinate system; the recalculated predicted position is the theoretical position of the feature points in the image coordinate system calculated by the forward projection principle based on the pose sequence of the corrected hand-eye matrix and the observation dataset, used to reflect the effect of the corrected coordinate projection; the actual observed position is the actual pixel coordinate of the feature points in the image sequence in the two-dimensional trajectory flow, obtained by the optical flow tracking algorithm, used as the benchmark for calculating the reprojection error; Euclidean distance is a mathematical method for measuring the spatial distance between two points, used to quantify the deviation between the recalculated predicted position and the actual observed position, realizing the accurate quantification of the reprojection error.

[0097] Specifically, firstly, based on the hand-eye matrix correction, the 6-dimensional Lie algebra vector is transformed into a transformation matrix through Lie algebra integration. The corrected hand-eye matrix is ​​obtained by multiplying it with the initial hand-eye matrix; then, based on the corrected hand-eye matrix and the pose sequence of the observation dataset, for each feature point, its position at each time sampling point is calculated using the forward projection principle. The recalculated predicted position, the forward projection principle is a mathematical method that projects three-dimensional spatial points onto a two-dimensional image plane based on the camera imaging model, and the calculation formula is: ,in This is the corrected hand-eye matrix. The coordinates of the feature points are given by the three-dimensional space coordinates. Finally, based on the recalculated predicted positions and the actual observed positions in the two-dimensional trajectory flow, the reprojection error of each feature point at each time sampling point is calculated using the Euclidean distance formula. The final reprojection error of each feature point is obtained by averaging all time sampling points.

[0098] This embodiment ensures the accuracy of the corrected hand-eye matrix based on Lie algebra integration, accurately calculates the re-predicted position of feature points through the forward projection principle, and uses Euclidean distance to quantify the deviation between the predicted position and the actual observed position, thus achieving high-precision calculation of reprojection error, improving the accuracy of error analysis, and providing reliable data support for subsequent error source separation and targeted compensation.

[0099] Furthermore, this embodiment provides a step for obtaining a set of rigid feature points and a set of reflective feature points based on reprojection error using an unsupervised clustering algorithm, including:

[0100] Based on the reprojection error and spatial coordinates of all feature points, preliminary clustering results are obtained by performing preliminary clustering using an unsupervised clustering algorithm;

[0101] Based on the preliminary clustering results, the optimized clustering boundaries are obtained through a region growing algorithm.

[0102] Based on the optimized clustering results, the feature points are divided into a set of rigid feature points with continuous spatial distribution and a set of reflective feature points with discrete spatial distribution through statistical analysis.

[0103] Among them, the unsupervised clustering algorithm is a density clustering algorithm, which includes a grouping process based on the reprojection error of feature points and the similarity of spatial coordinates, used for preliminary identification of error patterns; the region growing algorithm is an image processing algorithm based on spatial connectivity to expand the clustering region, using points in the initial cluster as seed points, and expanding the clustering region through the spatial distance threshold of adjacent points, used for merging adjacent clusters and optimizing cluster boundaries; the statistical analysis method is a method for quantitative evaluation based on the spatial distribution characteristics of clustering results, which determines the distribution type by calculating the spatial distribution dispersion of points in the cluster, used for accurate classification of rigid and reflective feature point sets; the rigid feature point set is a set of feature points with continuous spatial distribution and small dispersion, corresponding to geometric errors such as vehicle body flexible deformation, used for global hand-eye matrix update; the reflective feature point set is a set of feature points with discrete spatial distribution and large dispersion, corresponding to optical interference errors such as instrument reflection, used for local pixel compensation.

[0104] Specifically, firstly, based on the reprojection error of all feature points and their spatial coordinates in the image coordinate system, an unsupervised clustering algorithm is used to set the neighborhood radius. Preliminary clustering is performed using the minimum sample number parameter to obtain preliminary clustering results containing multiple clusters. Then, based on the preliminary clustering results, a region growing algorithm is used to start from the core point of each cluster and, according to a spatial distance threshold, proceed to the next cluster. Iterative merging of adjacent feature points yields optimized cluster boundaries; finally, based on the optimized clustering results, statistical analysis methods are used to calculate the spatial distribution dispersion of each cluster. and average neighborhood density This will satisfy the spatial distribution dispersion. Less than or equal to the preset dispersion threshold And average neighborhood density Greater than or equal to the preset neighborhood density threshold The cluster is determined to be a rigid feature point set, which satisfies the spatial distribution dispersion. Greater than the preset dispersion threshold or average neighborhood density Less than the preset neighborhood density threshold The cluster was determined to be a set of reflective feature points.

[0105] Based on reprojection error and spatial distribution characteristics, this embodiment achieves preliminary accurate division of feature point spatial distribution through unsupervised clustering, optimizes cluster boundaries and eliminates the influence of local isolated points through region growing algorithm, and achieves accurate classification of feature points by combining statistical analysis. This realizes effective separation of geometric deformation error and optical interference error, improves the reliability of error source identification and the pertinence of subsequent compensation.

[0106] Furthermore, this embodiment provides a step for constructing a local pixel compensation field based on a set of reflective feature points and their reprojection errors using a spatial interpolation algorithm, including:

[0107] Based on the spatial coordinates and reprojection error of the reflective feature point set, the pixel offset of each feature point is calculated by weighted averaging.

[0108] Based on the pixel offset and spatial distribution of all reflective feature points, an initial compensation field is established using a radial basis function interpolation algorithm.

[0109] Based on the requirement of continuity of the compensation field, the smoothness of the compensation field is optimized by thin plate spline spatial interpolation algorithm to obtain the final local pixel compensation field.

[0110] Among them, the pixel offset is the difference between the actual observed position and the predicted position of the reflective feature point, which is calculated by weighted averaging and is used to quantify the pixel-level deviation caused by local optical interference; the radial basis function interpolation algorithm is a mathematical method for constructing a continuous mapping based on the radial basis function, including the process of establishing an initial two-dimensional offset field based on the pixel offset and spatial coordinates of the reflective feature point, which is used to initially describe the local compensation relationship and establish a preliminary discrete compensation field; the thin plate spline spatial interpolation algorithm is a smooth interpolation method that preserves data point constraints and minimizes bending energy. It optimizes the spatial continuity of the initial compensation field by minimizing bending energy, and is used to obtain a continuous and natural local pixel compensation field.

[0111] Specifically, firstly, based on the spatial coordinates and reprojection error of the reflective feature point set, for each reflective feature point, its pixel offset is calculated by weighted averaging, using the following formula: ,in For reprojection error, The weights are based on spatial distance; then, based on the pixel offsets of all reflective feature points and their spatial distribution in the image coordinate system, an initial compensation field is constructed using a radial basis function interpolation algorithm, where the radial basis functions are... ,in Let Euclidean distance be the initial compensation field. Finally, based on the continuity requirement of the compensation field, the smoothness of the initial compensation field is optimized using a thin-plate spline spatial interpolation algorithm. The thin-plate spline algorithm obtains the final local pixel compensation field by minimizing the bending energy function, expressed as follows: ,in These are the weighting coefficients for the linear terms.

[0112] This embodiment calculates pixel offset based on reprojection error weighting to ensure the reliability of the offset. An initial compensation field is established through radial basis function interpolation to accurately reflect the offset pattern of the reflective area. The smoothness of the compensation field is optimized through thin plate spline interpolation, realizing the construction of a continuous and accurate local pixel compensation field. This effectively compensates for positional errors caused by optical interference such as instrument reflection, and improves the accuracy of image coordinate correction and the accuracy and stability of visual guidance.

[0113] Furthermore, this embodiment provides a step for obtaining the compensated assembly target coordinates through 3D reconstruction and coordinate compensation based on the updated hand-eye matrix and local pixel compensation field. The assembly target coordinates include vehicle body assembly coordinates and instrument installation coordinates. The steps include:

[0114] Based on real-time acquired binocular images, the initial pixel coordinates of vehicle body feature points and instrument panel feature points are obtained through feature extraction algorithms;

[0115] Based on the initial pixel coordinates of the instrument feature points, a local pixel compensation field is applied, and the compensated pixel coordinates of the instrument feature points are obtained through coordinate mapping.

[0116] Based on the compensated pixel coordinates of the instrument feature points, the instrument installation coordinates are calculated using the binocular triangulation principle.

[0117] Based on the initial pixel coordinates of the vehicle body feature points and the updated hand-eye matrix, the vehicle body assembly coordinates are calculated using the binocular triangulation principle.

[0118] The real-time acquired binocular images are the left and right eye images synchronously acquired by the binocular vision system at the current moment, used to provide the latest visual information of the workpiece; the feature extraction algorithm is the Harris corner detection algorithm, used to extract feature points of the vehicle assembly area and instrument body from the real-time acquired binocular images; the initial pixel coordinates are the original pixel positions of the feature points in the image coordinate system, obtained through the feature extraction algorithm; the local pixel compensation field is a two-dimensional pixel offset mapping function constructed based on the offset of the reflected feature points, used to compensate for optical interference errors in the instrument area; the compensated instrument feature point pixel coordinates are the instrument feature point pixel coordinates corrected by the local pixel compensation field, obtained through coordinate mapping. The binocular triangulation principle is a geometric method that uses the disparity relationship between the left and right eye images of a binocular vision system to calculate three-dimensional spatial coordinates, used to convert two-dimensional pixel coordinates into three-dimensional spatial coordinates; the updated hand-eye matrix is ​​a transformation matrix between the camera and the robotic arm tool coordinate system optimized by a rigid feature point set, used to provide accurate coordinate transformation relationships; the instrument mounting coordinates are the three-dimensional spatial coordinates of the instrument in the robot's base coordinate system, calculated using the compensated pixel coordinates of the instrument feature points and the binocular triangulation principle; the vehicle assembly coordinates are the three-dimensional spatial coordinates of the vehicle assembly holes in the robot's base coordinate system, calculated using the initial pixel coordinates of the vehicle feature points, the updated hand-eye matrix, and the binocular triangulation principle.

[0119] Specifically, such as Figure 3 As shown, firstly, based on real-time acquired binocular images, the Harris corner detection algorithm is used to extract feature points of the vehicle assembly area and instrument panel body from the left and right images respectively, resulting in a total of [number missing]. The initial pixel coordinates of each feature point, including Vehicle body feature points Instrument feature points Then, based on the initial pixel coordinates of the instrument feature points, the coordinate mapping function of the local pixel compensation field is used. The compensated pixel coordinates of each instrument feature point are calculated, i.e. , ,get The compensated pixel coordinates of the instrument feature points are then used. Based on these compensated pixel coordinates, the instrument installation coordinates are calculated using the binocular triangulation principle. The coordinates in the camera coordinate system are then transformed to the robot base coordinate system using the updated hand-eye matrix. Finally, based on the initial pixel coordinates of the vehicle body feature points and the updated hand-eye matrix, the vehicle body assembly coordinates are calculated using the binocular triangulation principle. The compensated assembly target coordinates are then obtained based on the instrument installation coordinates and the vehicle body assembly coordinates.

[0120] For example, this embodiment assumes that the real-time acquired binocular images contain Individual vehicle body feature points and Each instrument feature point is compensated for by local pixel field pairs. Coordinate compensation was performed on each instrument feature point, and the instrument installation coordinates were calculated using the binocular triangulation principle and the updated hand-eye matrix. and vehicle body assembly coordinates The instrument installation coordinates include A three-dimensional spatial point, the vehicle assembly coordinates include A three-dimensional point.

[0121] This embodiment accurately extracts the initial pixel coordinates of real-time feature points using the Harris corner detection algorithm, corrects pixel offset errors caused by instrument reflections through a local pixel compensation field, achieves accurate reconstruction of two-dimensional pixel coordinates to three-dimensional spatial coordinates through binocular triangulation, and ensures coordinate transformation accuracy through an updated hand-eye matrix. This enables dynamic compensation for vehicle body flexible deformation and instrument reflection interference, improving the positional accuracy and assembly quality stability of multi-robotic arm collaborative assembly.

[0122] Furthermore, embodiments of this application provide a dynamic vision guidance system for collaborative assembly of multiple robotic arms. This system includes a data acquisition module, a vision analysis module, an error compensation module, and a coordinate generation module.

[0123] Data acquisition module: Based on the initial hand-eye matrix, the assembly robot arm and the tightening robot arm are controlled to move collaboratively from their respective initial estimated poses to the corresponding reference target points. The image sequence is acquired synchronously through the binocular vision system, and the pose sequence of each robot arm is recorded synchronously through the robot controller to obtain the observation dataset.

[0124] Visual analysis module: Based on the observation dataset, the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained through optical flow tracking algorithm. The predicted motion trajectory of each feature point is obtained through forward projection principle, and the hand-eye matrix correction amount is obtained through nonlinear optimization algorithm.

[0125] Error compensation module: Based on the hand-eye matrix correction and the two-dimensional trajectory flow, the reprojection error of each feature point is obtained. Based on the reprojection error, the updated hand-eye matrix and local pixel compensation field are obtained through unsupervised clustering algorithm and spatial interpolation algorithm.

[0126] Coordinate generation module: Based on the updated hand-eye matrix and local pixel compensation field, the compensated assembly target coordinates are obtained through 3D reconstruction and coordinate compensation. The assembly target coordinates include the vehicle body assembly coordinates and the instrument installation coordinates.

[0127] The data acquisition module, visual analysis module, error compensation module, and coordinate generation module are all located on the server. The server receives data transmitted from the acquisition devices, including a binocular vision system and a robot controller, and performs further analysis. The data acquisition module is responsible for acquiring initial observation data and transmitting it to the visual analysis module. The visual analysis module extracts feature trajectories and calculates hand-eye matrix corrections based on the observation data, and transmits the results to the error compensation module. The error compensation module uses reprojection errors to separate error sources and generate compensation parameters, outputting an updated hand-eye matrix and local pixel compensation field to the coordinate generation module. Finally, the coordinate generation module outputs precise assembly target coordinates through 3D reconstruction and coordinate compensation to guide multi-robotic arm collaborative assembly. These modules are sequentially connected, forming a closed-loop process from data acquisition to coordinate generation, ensuring that the system dynamically and adaptively compensates for vehicle body deformation and instrument glare interference.

[0128] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0130] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0131] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A dynamic vision-guided method for collaborative assembly of multiple robotic arms, characterized in that, include: Based on the initial hand-eye matrix control, the assembly robot arm and the tightening robot arm move collaboratively from their respective initial estimated poses to the corresponding reference target points. The image sequence is synchronously acquired through the binocular vision system, and the pose sequence of each robot arm is synchronously recorded through the robot controller to obtain the observation dataset. Based on the observation dataset, the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained by optical flow tracking algorithm. The predicted motion trajectory of each feature point is obtained by forward projection principle, and the hand-eye matrix correction is obtained by nonlinear optimization algorithm. The reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow. Based on the reprojection error, the updated hand-eye matrix and local pixel compensation field are obtained through unsupervised clustering algorithm and spatial interpolation algorithm. Based on the updated hand-eye matrix and local pixel compensation field, the compensated assembly target coordinates are obtained through three-dimensional reconstruction and coordinate compensation. The assembly target coordinates include the vehicle body assembly coordinates and the instrument installation coordinates.

2. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 1, characterized in that, Based on the observed dataset, a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument panel is obtained through an optical flow tracing algorithm. The predicted motion trajectory of each feature point is obtained through forward projection, and the hand-eye matrix correction is obtained through a nonlinear optimization algorithm, including: Based on the image sequence of the observation dataset, a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body is obtained through an optical flow tracing algorithm; Based on the pose sequence, two-dimensional trajectory flow, and initial hand-eye matrix of the observation dataset, the predicted motion trajectory of each feature point is obtained through the forward projection principle, and the correction amount of the hand-eye matrix is ​​obtained through a nonlinear optimization algorithm.

3. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 2, characterized in that, The image sequence based on the observed dataset is used to obtain a two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument panel body through an optical flow tracing algorithm, including: Based on the image sequence of the observation dataset, the initial feature points of the vehicle body assembly area and the instrument body area are extracted; Based on the initial feature points, the motion trajectory of the feature points is obtained by tracking the feature points in consecutive image frames using an optical flow tracing algorithm. Based on the feature point matching relationship between the left and right eye images, a trajectory filtering algorithm is used to remove mismatched points, resulting in a stable two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body.

4. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 2, characterized in that, The pose sequence, two-dimensional trajectory flow, and initial hand-eye matrix based on the observed dataset are used to obtain the predicted motion trajectory of each feature point through the forward projection principle, and the hand-eye matrix correction is obtained through a nonlinear optimization algorithm, including: For each feature point in the two-dimensional trajectory flow, based on the pose sequence of the observation dataset and the initial hand-eye matrix, the predicted motion trajectory of each feature point is obtained through the forward projection principle. Based on the predicted motion trajectory of each feature point, the hand-eye matrix correction is obtained through a nonlinear optimization algorithm.

5. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 1, characterized in that, The process of obtaining the reprojection error of each feature point based on the hand-eye matrix correction and the two-dimensional trajectory flow, and then obtaining the updated hand-eye matrix and local pixel compensation field based on the reprojection error using an unsupervised clustering algorithm and a spatial interpolation algorithm, includes: The reprojection error of each feature point is obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow. Based on the reprojection error, an unsupervised clustering algorithm is used to obtain a set of rigid feature points and a set of reflective feature points; The updated hand-eye matrix is ​​obtained based on the rigid feature point set and its reprojection error. Based on the set of reflective feature points and their reprojection errors, a local pixel compensation field is constructed using a spatial interpolation algorithm.

6. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 5, characterized in that, The reprojection error of each feature point obtained based on the hand-eye matrix correction and the two-dimensional trajectory flow includes: Based on the aforementioned hand-eye matrix correction, the corrected hand-eye matrix is ​​obtained through Lie algebra integration. Based on the corrected hand-eye matrix and the pose sequence of the observation dataset, the recalculated predicted position of each feature point is obtained through the forward projection principle. Based on the recalculated predicted position and the actual observed position in the two-dimensional trajectory flow, the reprojection error of each feature point is calculated using Euclidean distance.

7. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 5, characterized in that, The process of obtaining the rigid feature point set and the reflective feature point set based on the reprojection error using an unsupervised clustering algorithm includes: Based on the reprojection error and spatial coordinates of all feature points, preliminary clustering results are obtained by performing preliminary clustering using an unsupervised clustering algorithm; Based on the preliminary clustering results, the optimized clustering boundaries are obtained using a region growing algorithm. Based on the optimized clustering results, the feature points are divided into a set of rigid feature points with continuous spatial distribution and a set of reflective feature points with discrete spatial distribution using statistical analysis methods.

8. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 5, characterized in that, The step of constructing a local pixel compensation field based on the set of reflective feature points and their reprojection errors using a spatial interpolation algorithm includes: Based on the spatial coordinates and reprojection error of the reflective feature point set, the pixel offset of each feature point is calculated by weighted averaging. Based on the pixel offset and spatial distribution of all reflective feature points, an initial compensation field is established using a radial basis function interpolation algorithm. Based on the requirement of continuity of the compensation field, the smoothness of the compensation field is optimized by thin plate spline spatial interpolation algorithm to obtain the final local pixel compensation field.

9. The dynamic vision guidance method for multi-robotic arm collaborative assembly according to claim 1, characterized in that, Based on the updated hand-eye matrix and local pixel compensation field, the compensated assembly target coordinates are obtained through 3D reconstruction and coordinate compensation. The assembly target coordinates include vehicle body assembly coordinates and instrument panel mounting coordinates, including: Based on real-time acquired binocular images, the initial pixel coordinates of vehicle body feature points and instrument panel feature points are obtained through feature extraction algorithms; Based on the initial pixel coordinates of the instrument feature points, a local pixel compensation field is applied, and the compensated pixel coordinates of the instrument feature points are obtained through coordinate mapping. Based on the compensated pixel coordinates of the instrument feature points, the instrument installation coordinates are calculated using the binocular triangulation principle. Based on the initial pixel coordinates of the vehicle body feature points and the updated hand-eye matrix, the vehicle body assembly coordinates are calculated using the binocular triangulation principle.

10. A dynamic vision guidance system for multi-robotic arm collaborative assembly, used to implement the dynamic vision guidance method for multi-robotic arm collaborative assembly as described in any one of claims 1-9, characterized in that, The dynamic vision guidance system for multi-robotic arm collaborative assembly includes: The data acquisition module controls the assembly robot arm and the tightening robot arm to move collaboratively from their respective initial estimated poses to the corresponding reference target points based on the initial hand-eye matrix. The module synchronously acquires image sequences through a binocular vision system and synchronously records the pose sequence of each robot arm through the robot controller to obtain the observation dataset. The visual analysis module, based on the observation dataset, obtains the two-dimensional trajectory flow of feature points in the vehicle assembly area and feature points in the instrument body through an optical flow tracking algorithm, obtains the predicted motion trajectory of each feature point through the forward projection principle, and obtains the hand-eye matrix correction amount through a nonlinear optimization algorithm. The error compensation module obtains the reprojection error of each feature point based on the hand-eye matrix correction and the two-dimensional trajectory flow, and obtains the updated hand-eye matrix and local pixel compensation field based on the reprojection error through unsupervised clustering algorithm and spatial interpolation algorithm. The coordinate generation module, based on the updated hand-eye matrix and local pixel compensation field, obtains the compensated assembly target coordinates through three-dimensional reconstruction and coordinate compensation. The assembly target coordinates include vehicle body assembly coordinates and instrument installation coordinates.

Citation Information

Patent Citations

  • Mechanical arm real-time tracking method based on binocular vision guidance

    CN112132894A

  • Industrial robot machining system for narrow space component and control method of industrial robot machining system

    CN118721244A