Template matching assembly positioning method, system and device with local pose invariance
Through the template matching method with constant local attitude, combined with the Sobel operator and RANSAC algorithm, the problems of insufficient accuracy and low efficiency in traditional robot assembly technology are solved, and high-precision and high-efficiency workpiece positioning and assembly are achieved, which is suitable for the flexible production of complex workpieces.
Patent Information
- Application Number
- CN202510552266.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Traditional robot assembly technology has shortcomings in high precision, high efficiency and flexible production, especially the global feature matching method is sensitive to noise, has high computational complexity, and weak generalization ability, resulting in poor positioning accuracy and real-time performance, making it difficult to adapt to the assembly needs of complex workpieces.
The template matching method with local pose unchanged is adopted, and the image edge is detected through the Sobel operator, and the key points are extracted in combination with the scale invariant feature transformation algorithm, the feature points are matched using the RANSAC algorithm, and the calibration information of the suction cup holder is combined to calculate the end position of the robot to achieve high-precision assembly.
It improves the positioning accuracy and robustness of robot assembly, reduces errors, improves assembly efficiency and adaptability, and is suitable for the flexible production of multiple types of workpieces.
Smart Images

Figure CN120070454B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot vision-guided assembly, and particularly relates to a template matching assembly positioning method, system and device with invariant local pose. Background Art
[0002] Robot assembly is gradually playing an increasingly important role in high-precision operations such as product manufacturing, precision electronic product processing, and aerospace component assembly. With the development of industrial automation and intelligent manufacturing technologies, traditional rigid assembly methods are difficult to meet the requirements of flexible and customized production. The robot assembly technology based on vision guidance has become one of the key technologies in modern manufacturing due to its high precision, flexibility and self-adaptive ability. Traditional robot assembly relies heavily on predefined paths and mechanical fixtures, which limits the flexibility and adaptability when dealing with complex or variable assembly tasks. In contrast, vision-guided assembly can achieve real-time perception and adaptive control, significantly improving precision, efficiency and versatility, providing an effective solution for automating the assembly process and improving the assembly quality. However, in the process of vision-guided automated assembly, errors may be introduced by sensors, robot joints and end effector tools, and the accumulation of these errors still poses a challenge to the high-precision implementation of robot assembly.
[0003] The core challenge of robot vision-guided assembly technology lies in how to achieve high-precision, high-efficiency and robust workpiece positioning. Existing technologies mainly rely on global feature matching or contact detection methods, but still have the following significant defects:
[0004] 1. Insufficient precision and error superposition
[0005] Traditional global template matching methods (such as normalized cross-correlation matching, edge contour matching) are sensitive to image noise, local occlusion and illumination changes, and it is difficult to achieve high positioning precision. The superposition of the robot's own repeated positioning error and the absolute positioning error of the vision system (caused by camera calibration deviation, image distortion, etc.) further reduces the overall accuracy of the system.
[0006] 2. Low efficiency and poor real-time performance
[0007] The global search algorithm needs to traverse the entire image, with high computational complexity. Taking an image with a resolution of 1280×1024 as an example, the full-image edge matching takes more than 200 ms, which is difficult to meet the real-time requirement (usually required ≤50 ms) of a high-speed production line. Contact detection (such as mechanical probe positioning) requires physical contact with the workpiece, increasing the single detection time by 30~50 ms and may scratch the surface of the precision workpiece.
[0008] 3. Weak generalization ability
[0009] The existing methods have poor adaptability to changes in the shape, pose, and distance of workpieces. Methods based on single contour matching cannot handle workpiece rotation or scaling (e.g., the matching failure rate > 60% when the rotation exceeds ±30°); methods based on color segmentation fail under light fluctuations or background interference (e.g., a ±20% change in light intensity leads to a 40% increase in the false detection rate). When switching between workpieces of multiple varieties, it is necessary to redesign the positioning algorithm, resulting in high transformation costs for flexible production lines.
[0010] Based on this, a template matching assembly positioning method, system, and device with local pose invariance are proposed. Summary of the Invention
[0011] Aiming at the above technical problems, the present invention provides a template matching assembly positioning method, system, and device with local pose invariance.
[0012] The technical solution adopted by the present invention to solve its technical problems is:
[0013] A template matching assembly positioning method with local pose invariance, the method includes the following steps:
[0014] S100: Collect image data of the workpiece to be assembled and the assembly environment;
[0015] S200: Apply the Sobel operator to perform global rough positioning on the image data, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate area of the workpiece;
[0016] S300: In the candidate area, use the scale-invariant feature transform algorithm to extract the key points of the workpiece, and generate local feature descriptors based on the key points;
[0017] S400: Match the local feature descriptors with the features of the pre-stored template, and estimate the pose information of the workpiece through the matched feature points;
[0018] S500: Based on the constraint conditions of the fixed points of the suction cup for clamping the workpiece, combined with the transformation matrix from the pre-calibrated coordinate system C of the suction nozzle vacuum gripper to the workpiece coordinate system W, solve the assembly pose of the robot end, and control the robot to complete the assembly operation of the workpiece according to the assembly pose.
[0019] Preferably, the Sobel operator in S200 includes a horizontal edge detection convolution kernel and a vertical edge detection convolution kernel, which are respectively used to detect the gradient changes in the horizontal and vertical directions of the image. S200 includes:
[0020] S210: For each pixel point, by performing convolution operations on the input image with the two convolution kernels in the Sobel operator, obtain the gradient images of the image in the horizontal and vertical directions and ;
[0021] S220: Calculate the gradient magnitude M and the gradient direction based on the gradient images of the image in the horizontal and vertical directions and calculate the gradient magnitude M and the gradient direction, specifically: , specifically as follows:
[0022] ;
[0023] S230: Consider all points with a gradient magnitude less than the preset threshold T as the background, and consider all points with a gradient magnitude greater than the preset threshold T as the edges, extract the edges in the image, and represent the edge information with a binary image, where the edge part is white and the other parts are black.
[0024] Preferably, S300 includes:
[0025] S310: Construct a Gaussian pyramid by continuously blurring and downsampling the image to generate a multi-scale image representation;
[0026] S320: Detect scale-space extreme points on the images at different scales through the Gaussian pyramid, so as to find the key points with local scale invariance in the image;
[0027] S330: In order to make the key points with scale invariance have rotational invariance, determine the main direction of each key point based on the gradient direction information of the local area of the image;
[0028] S340: Take the key point as the center, divide the local area into multiple small blocks. For each small block, calculate the direction and magnitude of the gradient, construct an orientation histogram, merge the gradient orientation histograms of all small blocks, generate a descriptor containing 128 elements, and finally normalize the 128-dimensional descriptor to make it have a unit length.
[0029] Preferably, S310 is specifically: Apply a Gaussian filter to blur the image:
[0030] ;
[0031] where is the image after Gaussian blur, represents the scale of the current layer, and respectively represent the coordinates of the input image;
[0032] Downsample the blurred image by a certain ratio to generate the image of the next layer;
[0033] Repeat the image blurring and downsampling process, and build a new layer on the basis of the image of the previous layer each time. Finally, the obtained Gaussian pyramid contains images of multiple scales;
[0034] S320 includes:
[0035] Generate a Difference-of-Gaussians (DoG) image by subtracting Gaussian images at adjacent scales, specifically:
[0036] ;
[0037] where k is the scale increase factor, is the Difference-of-Gaussians (DoG) image;
[0038] Take a pixel in the image at the current scale and compare it with the pixels of its 26 adjacent points. If the pixel value of the current pixel point is greater than or less than the pixel values of all 26 neighborhood points, it is a local maximum point or a local minimum point.
[0039] Preferably, S330 includes:
[0040] S331: Calculate the gradient magnitude and direction in the local region on the scale image where the key point is located, specifically:
[0041] ;
[0042] ;
[0043] where, and are the gradients of the image in the x and y directions respectively, θ is the gradient direction, and M is the gradient magnitude;
[0044] S332: Calculate the weighted sum of the gradient magnitude and gradient direction in each direction interval of the key point, and select the direction with the largest peak in the histogram as the main direction of the key point.
[0045] Preferably, S400 includes:
[0046] S410: Calculate the Euclidean distance between the local feature descriptor in the image to be matched and the feature descriptor of the pre-stored template image, specifically:
[0047] ;
[0048] where, and are the values of the i-th element in the local feature descriptor D1 and the feature descriptor D2 of the pre-stored template image, and n is the dimension in the descriptor;
[0049] Select the descriptor pair with the smallest distance as the matching point;
[0050] S420: Using the RANSAC algorithm, randomly select a preset number of pairs of matching points to estimate the homography transformation matrix, and obtain the estimated homography transformation matrix , and use the estimated homography transformation matrix to determine whether other matching points are consistent with the current transformation; repeat random sampling, and finally obtain a largest inlier set, thereby effectively removing mis-matching points;
[0051] S430: Estimate the pose of the object by calculating the homography transformation matrix between the matching points in the largest inlier set to minimize the error of the matching points.
[0052] Preferably, S420 includes:
[0053] Using the RANSAC algorithm, randomly select 4 pairs of matching points from all sets of matching points to form a sampling set for estimating the homography transformation matrix, and obtain the estimated homography transformation matrix , specifically:
[0054] ;
[0055] Calculate the reprojection error of all matching points through the estimated homography transformation matrix :
[0056] ;
[0057] wherein, is the original image coordinate actually detected, is the target image coordinate that should be mapped theoretically according to the estimated homography transformation matrix;
[0058] If the reprojection error is less than the set threshold, the matching point is an inlier, otherwise it is an outlier. Repeat the above process, finally determine the largest inlier set, and remove mis-matching points;
[0059] S430 includes:
[0060] Use all matching points in the largest inlier set to calculate the homography transformation matrix , specifically:
[0061] ;
[0062] wherein, is the point in the original image of the j-th pair of matching points , is the transformed point of the j-th pair of matching points , and m is the total number of matching points;
[0063] wherein, for the matching point pair {( , )} and {( , )}, the homography transformation matrix satisfies the following relationship:
[0064] ;
[0065] is in the form of:
[0066] ;
[0067] Among them, the homography transformation matrix describes the transformation under different perspective views on the same plane, , , , control rotation, scaling, and shear transformations; , control translation transformation, , control perspective distortion, is the normalization factor;
[0068] Decompose the homography transformation matrix to obtain the rotation matrix R and the translation vector T:
[0069] ;
[0070] Restore the pose of the object in three-dimensional space through the rotation matrix R and the translation vector T obtained by decomposition.
[0071] Preferably, S500 includes:
[0072] S510: Define the pose of the vacuum gripper at the end of the robot in the base coordinate system B, where the pose of the vacuum gripper at the end includes position and orientation;
[0073] S520: Calibrate the transformation matrix from the coordinate system C of the suction nozzle vacuum gripper to the workpiece coordinate system W, that is, transform the feature points from the gripper coordinate system C to the workpiece coordinate system W;
[0074] S530: Determine the pose of the workpiece in the gripper coordinate system through the position and orientation information of the vacuum gripper at the end of the robot, combined with the fixed constraint conditions between the suction nozzle vacuum gripper and the workpiece. The cooperation between the suction nozzle and the workpiece needs to meet the double coordinate system alignment condition:
[0075] ;
[0076] Among them, is the fixed mating pose, Indicates the expected pose of the gripper coordinate system relative to the workpiece coordinate system when the nozzle contacts the workpiece, which is the expected coordinate of the center origin of the gripper in the workpiece coordinate system when the nozzle contacts the workpiece;
[0077] S540: According to the transformation matrix from the nozzle vacuum gripper coordinate system C to the workpiece coordinate system W and combined with the pose of the workpiece in the base coordinate system, calculate the actual pose of the nozzle vacuum gripper at the end of the robot in the base coordinate system:
[0078] ;
[0079] where, is the pose of the gripper in the base coordinate system, is the pose of the workpiece in the base coordinate system, is the transformation matrix from the nozzle vacuum gripper coordinate system to the workpiece coordinate system;
[0080] S550: According to the calculated pose information of the end effector, the robot adjusts its motion trajectory to ensure that the gripper is correctly fixed at the specified position of the workpiece, and performs correction in combination with the feedback of the machine vision system, so as to ensure the accurate assembly pose.
[0081] The template matching assembly positioning system with local pose invariance includes an image data acquisition module, a candidate region determination module, a local feature descriptor generation module, a workpiece pose information estimation module, and a workpiece assembly module;
[0082] The image data acquisition module is used to acquire the image data of the workpiece to be assembled and the assembly environment;
[0083] The candidate region determination module is used to perform global rough positioning on the image data by applying the Sobel operator, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate region of the workpiece;
[0084] The local feature descriptor generation module is used to extract the key points of the workpiece in the candidate region by using the scale-invariant feature transform algorithm, and generate local feature descriptors based on the key points;
[0085] The workpiece pose information estimation module is used to match the local feature descriptors with the features of the pre-stored template, and estimate the pose information of the workpiece through the matched feature points;
[0086] The workpiece assembly module is used to solve the assembly pose of the end of the robot based on the constraint conditions of the fixed points of the suction cup for gripping the workpiece, combined with the pre-calibrated transformation matrix from the nozzle vacuum gripper coordinate system C to the workpiece coordinate system W, and control the robot to complete the assembly operation of the workpiece according to the assembly pose.
[0087] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a template matching assembly positioning method with local pose invariance are implemented.
[0088] The above-mentioned template matching assembly positioning method, system, and device with local pose invariance are an advanced method that combines precise image processing, feature matching, and a machine vision system with robot control, aiming to improve the accuracy and efficiency of robots during the assembly process. Its advantages lie in high-precision positioning, strong robustness, multi-step processing optimization, dynamic adaptability, and the close cooperation between the robot and the vision system. Technologies such as the RANSAC algorithm and suction cup gripper calibration are used to ensure the precise identification and assembly of workpieces, reduce manual intervention, are widely applicable to high-precision assembly tasks, and have significant application potential. Brief Description of the Drawings
[0089] Figure 1 It is a flowchart of a template matching assembly positioning method with local pose invariance in an embodiment of the present invention;
[0090] Figure 2 It is a schematic diagram of obtaining robot image data provided in an embodiment of the present invention. Detailed Embodiments
[0091] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0092] In one embodiment, as Figure 1 shown, a template matching assembly positioning method with local pose invariance, the method includes the following steps:
[0093] S100: Collect image data of the workpiece to be assembled and the assembly environment;
[0094] S200: Apply the Sobel operator to perform global rough positioning on the image data, detect the edge information in the image through the gradient changes of horizontal and vertical edges, and determine the candidate area of the workpiece;
[0095] S300: In the candidate area, use the scale-invariant feature transform algorithm to extract the key points of the workpiece, and generate local feature descriptors based on the key points;
[0096] S400: Match the local feature descriptors with the features of the pre-stored template, and estimate the pose information of the workpiece through the matched feature points;
[0097] S500: Based on the constraint conditions of the workpiece fixing points clamped by the suction cup, combined with the pre-calibrated transformation matrix from the coordinate system C of the nozzle vacuum gripper to the workpiece coordinate system W, calculate the assembly pose of the robot end, and control the robot to complete the workpiece assembly operation according to the assembly pose.
[0098] Compared with other assembly positioning methods, the present invention can accurately capture and locate the key points of the workpiece through local feature extraction and template matching, greatly improving the positioning accuracy in terms of accuracy, robustness, and adaptability. Especially in complex and deformed workpiece environments, the use of local pose invariant descriptors can reduce errors caused by object pose changes or different perspectives, thus effectively overcoming the limitations of traditional methods and achieving precise robot assembly positioning.
[0099] Further, as Figure 2 shown in the figure, in the figure, B is the robot base coordinate system, C is the coordinate system of the nozzle vacuum gripper, W is the workpiece coordinate system, S is the camera. After the robot picks up the workpiece on the carrier plate, it moves to the lower camera photographing point to obtain an image. After determining the camera parameters using the calibration method, execute the above S100 - S500.
[0100] In one embodiment, the Sobel operator in S200 includes a horizontal edge detection convolution kernel and a vertical edge detection convolution kernel, which are respectively used to detect the gradient changes in the horizontal and vertical directions of the image. S200 includes:
[0101] S210: For each pixel point, by convolving the two convolution kernels in the Sobel operator with the input image (grayscale image), obtain the gradient images of the image in the horizontal and vertical directions and ;
[0102] S220: According to the gradient images of the image in the horizontal and vertical directions and calculate the gradient magnitude M (edge intensity) and the gradient direction , specifically:
[0103] ;
[0104] S230: In order to better highlight the edges in the image, usually threshold the gradient magnitude. Consider all points with gradient magnitudes less than the preset threshold T as the background (i.e., not belonging to the edge), and consider all points with gradient magnitudes greater than the preset threshold T as the edge, extract the edges in the image, and represent the edge information with a binary image, where the edge part is white and other parts are black.
[0105] In one embodiment, S300 includes:
[0106] S310: Construct a Gaussian pyramid by continuously blurring and downsampling the image to generate a multi-scale image representation;
[0107] S320: Detect scale-space extreme points on the images at different scales through the Gaussian pyramid, thereby finding the key points with local scale invariance in the image;
[0108] S330: To make the key points with scale invariance also have rotational invariance, determine the main direction of each key point based on the gradient direction information of the local region of the image;
[0109] S340: Centered on the key points, divide the local region into multiple small blocks. For each small block, calculate the direction and magnitude of the gradient, construct an orientation histogram, merge the gradient orientation histograms of all small blocks, generate a descriptor containing 128 elements, and finally normalize the 128-dimensional descriptor to make it have unit length.
[0110] In one embodiment, S310 is specifically: Apply a Gaussian filter to blur the image:
[0111] ;
[0112] where, is the image after Gaussian blur, represents the scale of the current layer, and respectively represent the coordinates of the input image;
[0113] Downsample the blurred image by a certain ratio to generate the image of the next layer;
[0114] Repeat the image blurring and downsampling process. Each time, construct a new layer based on the image of the previous layer. Finally, obtain the Gaussian pyramid, which contains images of multiple scales;
[0115] S320 includes:
[0116] Generate a Difference-of-Gaussians (DoG) image by subtracting adjacent-scale Gaussian images. Specifically:
[0117] ;
[0118] where k is the scale increase factor, is the Difference-of-Gaussians (DoG) image;
[0119] Take a pixel in the image of the current scale and compare it with the pixels of its adjacent 26 points. If the pixel value of the current pixel point is greater than or less than the pixel values of all 26 neighborhood points, it is a local maximum point or a local minimum point.
[0120] In one embodiment, S330 includes:
[0121] S331: On the scale image where the key point is located, calculate the gradient magnitude and direction of the local region, specifically:
[0122] ;
[0123] ;
[0124] Wherein, and are the gradients of the image in the x and y directions respectively, θ is the gradient direction, and M is the gradient magnitude;
[0125] S332: Calculate the weighted sum of the gradient magnitude and gradient direction within each direction interval of the key point, and select the direction with the largest peak in the histogram as the main direction of the key point.
[0126] In one embodiment, S400 includes:
[0127] S410: Calculate the Euclidean distance between the local feature descriptor in the image to be matched and the feature descriptor of the pre-stored template image, specifically:
[0128] ;
[0129] Wherein, and are the values of the i-th element in the local feature descriptor D1 and the feature descriptor D2 of the pre-stored template image, and n is the dimension in the descriptor;
[0130] Select the descriptor pair with the smallest distance as the matching point; generally speaking, the smaller the Euclidean distance of the descriptor, the more similar these two feature points are;
[0131] S420: Use the RANSAC algorithm to estimate the homography transformation matrix by randomly selecting a preset number of pairs of matching points, and obtain the estimated homography transformation matrix , and use the estimated homography transformation matrix to judge whether other matching points are consistent with the current transformation; repeat random sampling, and finally obtain a maximum inlier set, thereby effectively eliminating mis-matching points;
[0132] S430: Estimate the pose of the object by calculating the homography transformation matrix among the matching points in the maximum inlier set to minimize the error of the matching points.
[0133] In one embodiment, S420 includes:
[0134] Randomly select 4 pairs of matching points from all sets of matching points using the RANSAC algorithm to form a sampling set for estimating the homography transformation matrix and obtain the estimated homography transformation matrix Specifically
[0135] ;
[0136] Calculate the reprojection error of all matching points using the estimated homography transformation matrix :
[0137] ;
[0138] where is the original image coordinate actually detected and
[0139] is the target image coordinate that should be mapped theoretically according to the estimated homography transformation matrix;
[0140] S430 includes
[0141] Calculate the homography transformation matrix using all matching points in the maximum inlier set Specifically
[0142] ;
[0143] where is the point in the original image of the j-th pair of matching points , is the transformed point of the j-th pair of matching points and m is the total number of matching points;
[0144] where, for the matching point pairs {( , )} and {( , )}, the homography transformation matrix satisfies the following relationship
[0145] ;
[0146] The form of
[0147] ;
[0148] where the homography transformation matrix describes the transformation under different perspective views on the same plane , , , Control rotation, scaling, and shear transformations; , Control translation transformation, , Control perspective distortion, is a normalization factor, usually set to 1;
[0149] Decompose the homography transformation matrix to obtain the rotation matrix R and the translation vector T:
[0150] ;
[0151] Restore the pose of the object in 3D space through the rotation matrix R and the translation vector T obtained by decomposition.
[0152] In one embodiment, S500 includes:
[0153] S510: Define the pose of the robot end effector (suction cup vacuum gripper) in the base coordinate system, where the pose of the end effector includes position (x, y, z coordinates) and orientation (rotation angle or rotation matrix); usually, the robot control system uses a homogeneous transformation matrix to represent this pose;
[0154] S520: Calibrate the transformation matrix from the suction cup vacuum gripper coordinate system C to the workpiece coordinate system W; further, the purpose of calibrating the transformation matrix is to associate the pose of the suction cup gripper with the pose of the workpiece. The transformation matrix is usually obtained through the calibration process, which describes the position and orientation differences between the gripper coordinate system and the workpiece coordinate system. Through calibration, a homogeneous transformation matrix from the gripper coordinate system to the workpiece coordinate system can be obtained ;
[0155] S530: Determine the position of the gripper in the workpiece coordinate system based on the position and orientation information of the robot end effector and the fixed constraint conditions between the suction cup vacuum gripper and the workpiece; further, the relative position and orientation of the suction cup on the workpiece are determined, the pose of the robot when the suction cup picks up the workpiece to be assembled is determined, and the suction cup vacuum gripper is at the end of the robot. Therefore, the relative position between the workpiece and the suction cup is used as a fixed constraint condition, and the cooperation between the suction cup and the workpiece needs to meet the double coordinate system alignment condition:
[0156] ;
[0157] where, is the fixed mating pose, represents the expected pose of the gripper coordinate system relative to the workpiece coordinate system when the suction cup contacts the workpiece, is the expected coordinate of the center origin of the gripper in the workpiece coordinate system when the nozzle contacts the workpiece;
[0158] S540: According to the transformation matrix from the coordinate system C of the nozzle vacuum gripper to the workpiece coordinate system W , combined with the pose of the end effector of the robot in the base coordinate system, calculate the actual pose of the end effector of the robot in the workpiece coordinate system:
[0159] ;
[0160] wherein, is the pose of the gripper in the base coordinate system, is the pose of the workpiece in the base coordinate system, is the transformation matrix from the coordinate system of the nozzle vacuum gripper to the workpiece coordinate system;
[0161] S550: According to the calculated pose information of the end effector, the robot adjusts its motion trajectory to ensure that the gripper is correctly fixed at the specified position of the workpiece, and performs correction in combination with the feedback of the machine vision system, so as to ensure the precise assembly pose.
[0162] In a detailed embodiment, the present invention is tested on a SCARA robot, image data of the workpiece to be assembled and its environment are collected, global rough positioning is performed on the image data, the gradient information of the image is detected by the Sobel operator, and threshold processing is performed on the image based on the gradient amplitude and direction, so as to extract the edge information in the image. Key points of the image are extracted through the Gaussian pyramid, and combined with the extreme points in the scale space, key local feature points in the image can be detected. The main direction of each key point is calculated through the weighted histogram of the gradient direction, the extracted local pose invariant descriptor is matched with the features of the pre-stored template, and the Euclidean distance and the RANSAC algorithm are used to screen the optimal matching points. The RANSAC algorithm randomly selects matching points, estimates the homography transformation matrix, and thus eliminates the mismatched points. After obtaining the pose of the workpiece, combined with the calibration information of the suction cup gripper, the pose of the gripper is docked with the coordinate system of the workpiece. Combining the feedback of the robot control system and the machine vision system, ensure that the robot can perform precise assembly operations according to the calculated assembly pose, reduce human intervention, and achieve the high-precision assembly goal.
[0163] Table 1 Positioning offset
[0164]
[0165] The above-mentioned template matching positioning method based on local pose invariance includes: collecting image data of the workpiece to be assembled and the assembly environment; implementing an assembly strategy that combines global rough positioning and fine positioning with local pose invariants to achieve high-precision assembly of the robot under visual guidance. The local feature extraction algorithm is used to extract key points in the image. Assuming that the suction cup can always suck a fixed point on the workpiece each time, the transformation from the gripper to the workpiece coordinate system is always a fixed value. By calibrating the transformation matrix from the workpiece to the end coordinate, the assembly coordinate pose of the robot can be indirectly determined; the workpiece position is located by matching the information with local pose invariance. Through the matched feature points, the pose information such as the relative position, angle, and rotation of the target object is estimated. The present invention provides a new method for calculating the assembly pose by combining visual information with the local reference pose of the robot, solving the problem of inaccurate repetitive positioning accuracy of the robot body and absolute positioning accuracy of the system; the above technical solution overcomes the technical problems of low efficiency, poor real-time performance, and contact inspection in the existing planar positioning technology, effectively improving the detection efficiency and detection accuracy, and the method of the present invention can also solve the problem of poor generalization ability for workpieces with different shapes, poses, and distances.
[0166] A template matching assembly positioning system with local pose invariance includes an image data acquisition module, a candidate region determination module, a local feature descriptor generation module, a workpiece pose information estimation module, and a workpiece assembly module;
[0167] The image data acquisition module is used to collect image data of the workpiece to be assembled and the assembly environment;
[0168] The candidate region determination module is used to perform global rough positioning on the image data by applying the Sobel operator, detect the edge information in the image through the gradient changes of horizontal and vertical edges, and determine the candidate region of the workpiece;
[0169] The local feature descriptor generation module is used to extract the key points of the workpiece in the candidate region by using the scale-invariant feature transform algorithm and generate local feature descriptors based on the key points;
[0170] The workpiece pose information estimation module is used to match the local feature descriptors with the features of the pre-stored template and estimate the pose information of the workpiece through the matched feature points;
[0171] The workpiece assembly module is used to, based on the constraint condition of the suction cup gripping the fixed point of the workpiece, combine the pre-calibrated transformation matrix from the suction nozzle vacuum gripper coordinate system C to the workpiece coordinate system W, calculate the assembly pose of the robot end, and control the robot to complete the assembly operation of the workpiece according to the assembly pose.
[0172] For the specific limitations of a local pose-invariant template matching assembly positioning system, reference may be made to the limitations of a local pose-invariant template matching assembly positioning method in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned local pose-invariant template matching assembly positioning system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0173] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a local pose-invariant template matching assembly positioning method are implemented.
[0174] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0175] The above has introduced in detail a local pose-invariant template matching assembly positioning method, system, and device provided by the present invention. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A local pose-invariant template matching assembly positioning method, characterized in that The method includes the following steps: S100: Collect image data of the workpiece to be assembled and the assembly environment; S200: Apply the Sobel operator to perform global rough positioning on the image data, detect edge information in the image through the gradient changes of horizontal and vertical edges, and determine the candidate region of the workpiece; S300: In the candidate region, use the Scale-Invariant Feature Transform (SIFT) algorithm to extract key points of the workpiece, and generate local feature descriptors based on the key points; S400: Match the local feature descriptors with the features of the pre-stored template, and estimate the pose information of the workpiece through the matched key points; S500: Based on the fixed constraint conditions for clamping the workpiece by the vacuum gripper at the end of the robot, combined with the transformation matrix from the coordinate system C of the vacuum gripper at the end of the robot to the workpiece coordinate system W that has been pre-calibrated, solve the assembly pose of the vacuum gripper at the end of the robot, and control the robot to complete the assembly operation of the workpiece according to the assembly pose; S500 includes: S510: Define the pose of the vacuum gripper at the end of the robot in the base coordinate system B, where the pose of the end vacuum gripper includes position and orientation; S520: Calibrate the transformation matrix from the coordinate system C of the suction nozzle vacuum gripper to the workpiece coordinate system W ; S530: Through the position and orientation information of the vacuum gripper at the end of the robot, combined with the fixed constraint conditions between the vacuum gripper and the workpiece, determine the pose of the gripper in the workpiece coordinate system. The cooperation between the suction nozzle and the workpiece needs to meet the double coordinate system alignment condition: ; Among them, is the fixed mating pose, which represents the expected pose of the gripper coordinate system relative to the workpiece coordinate system when the nozzle contacts the workpiece, and is the expected coordinate of the center origin of the gripper in the workpiece coordinate system when the nozzle contacts the workpiece; S540: Based on the transformation matrix from the coordinate system C of the nozzle vacuum gripper to the workpiece coordinate system W , combined with the pose of the workpiece in the base coordinate system, calculate the actual pose of the nozzle vacuum gripper at the end of the robot in the base coordinate system: ; Among them, is the pose of the gripper in the base coordinate system, is the pose of the workpiece in the base coordinate system, is the transformation matrix from the coordinate system of the suction nozzle vacuum gripper to the workpiece coordinate system; S550: According to the calculated pose information of the vacuum gripper at the end of the robot, the robot adjusts its motion trajectory to ensure that the gripper is correctly fixed at the specified position of the workpiece, and performs correction in combination with the feedback of the machine vision system, so as to ensure an accurate assembly pose.
2. The method according to claim 1, wherein The Sobel operator in S200 includes a horizontal edge detection convolution kernel and a vertical edge detection convolution kernel, which are respectively used to detect the gradient changes in the horizontal and vertical directions of the image. S200 includes: S210: For each pixel, by convolving two convolution kernels in the Sobel operator with the input image, gradient images of the image in the horizontal and vertical directions are obtained and ; S220: According to the gradient images of the image in the horizontal and vertical directions and calculate the gradient magnitude M and the gradient direction , specifically as follows: ; S230: Consider all points with gradient magnitudes less than the preset threshold T as the background, and consider all points with gradient magnitudes greater than the preset threshold T as edges, extract the edges in the image, and represent the edge information with a binary image, where the edge part is white and the other parts are black.
3. The method according to claim 2, wherein S300 includes: S310: Construct a Gaussian pyramid by continuously blurring and downsampling the image to generate a multi-scale image representation; S320: Detect scale-space extreme points on the images at different scales through the Gaussian pyramid, so as to find key points with local scale invariance in the image; S330: In order to make the key points with scale invariance have rotation invariance, determine the main direction of each key point based on the gradient direction information of the local region of the image; S340: Centered on the key points, divide the local region into multiple small blocks. For each small block, calculate the direction and magnitude of the gradient, construct a direction histogram, merge the gradient direction histograms of all small blocks, generate a descriptor containing 128 elements, and finally normalize the 128-dimensional descriptor to make it have unit length.
4. The method according to claim 3, wherein S310 is specifically: Apply a Gaussian filter to perform Gaussian blur on the image: ; Among them, is the image after Gaussian blur, represents the scale of the current layer, and respectively represent the coordinates of the input image; Downsample the blurred image by a certain ratio to generate the image of the next layer; Repeat the image blurring and downsampling processes, each time constructing a new layer based on the image of the previous layer, and finally obtain the Gaussian pyramid, which contains images of multiple scales; S320 includes: Generate a Difference of Gaussian (DoG) image by subtracting adjacent-scale Gaussian images, specifically: ; where k is the scale increase factor, is the Difference of Gaussian image; Take a pixel in the image at the current scale and compare it with the pixels of its 26 adjacent points. If the pixel value of the current pixel point is greater than or less than the pixel values of all 26 neighborhood points, it is a local maximum point or a local minimum point.
5. The method according to claim 4, wherein S330 includes: S331: On the scale image where the key point is located, calculate the gradient magnitude and direction of the local region, specifically: ; ; wherein, and are the gradients of the image in the x and y directions respectively, θ is the gradient direction, and M is the gradient magnitude; S332: Calculate the weighted sum of the gradient magnitudes and gradient directions within each direction interval of the key point, and select the direction with the largest peak in the histogram as the main direction of the key point.
6. The method according to claim 5, characterized in that S400 includes: S410: Calculate the Euclidean distance between the local feature descriptor in the image to be matched and the feature descriptor of the pre-stored template image, specifically: ; Wherein, and are the values of the i-th elements in the local feature descriptor D1 and the feature descriptor D2 of the pre-stored template image, and n is the dimension in the descriptor; Select the descriptor pair with the smallest distance as the matching point; S420: Using the RANSAC algorithm, estimate the homography transformation matrix by randomly selecting a preset number of pairs of matching points, and obtain the estimated homography transformation matrix , and use the estimated homography transformation matrix to determine whether other matching points are consistent with the current transformation; repeat random sampling, and finally obtain a maximum inlier set, thereby effectively removing mismatched points; S430: Estimate the pose of the workpiece by calculating the homography transformation matrix between the matching points in the maximum inlier set to minimize the error of the matching points.
7. The method according to claim 6, wherein S420 includes: Randomly select 4 pairs of matching points from all sets of matching points using the RANSAC algorithm to form a sampling set for estimating the homography transformation matrix and obtain the estimated homography transformation matrix Specifically ; Calculate the reprojection error of all matching points using the estimated homography transformation matrix : ; Among them, is the original image coordinate actually detected, is the target image coordinate that should be mapped theoretically according to the estimated homography transformation matrix; If the reprojection error is less than the set threshold, then the matching point is an inlier, otherwise it is an outlier. Repeat the above process to finally determine the largest set of inliers and eliminate the mismatched points; S430 includes: Calculate the homography transformation matrix using all the matching points in the maximum inlier set , specifically: ; Among them, is the point in the original image among the j-th pair of matching points , is the transformed point among the j-th pair of matching points , and m is the total number of matching points; Among them, for the matching point pairs {( , )} and {( , ), the homography transformation matrix satisfies the following relationship: ; is in the form of: ; Among them, the homography transformation matrix describes the transformation under different perspective views on the same plane, , , , controls rotation, scaling, and shear transformations; , controls translation transformation, , controls perspective distortion, is the normalization factor; Decompose the homography transformation matrix Obtain the rotation matrix R and the translation vector T: ; Restore the pose of the workpiece in the three-dimensional space through the rotation matrix R and translation vector T obtained by decomposition.
8. A local pose-invariant template matching assembly positioning system using the method according to any one of claims 1 to 7, characterized in that, Includes an image data acquisition module, a candidate region determination module, a local feature descriptor generation module, a workpiece pose information estimation module, and a workpiece assembly module; The image data acquisition module is used to acquire the image data of the workpiece to be assembled and the assembly environment; The candidate region determination module is used to globally and roughly locate the image data by applying the Sobel operator, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate region of the workpiece; The local feature descriptor generation module is used to extract the key points of the workpiece within the candidate region by using the Scale-Invariant Feature Transform (SIFT) algorithm and generate local feature descriptors based on the key points; The workpiece pose information estimation module is used to match the local feature descriptors with the features of the pre-stored template and estimate the pose information of the workpiece through the matched key points; The workpiece assembly module is used to, based on the fixed constraint conditions for gripping the workpiece by the vacuum gripper at the end of the robot, combined with the transformation matrix from the pre-calibrated coordinate system C of the vacuum gripper at the end of the robot to the workpiece coordinate system W, solve the assembly pose of the vacuum gripper at the end of the robot and control the robot to complete the workpiece assembly operation according to the assembly pose.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A flexible assembly system and method based on multi-mode information description
CN109543823A
Visual guidance right-angle robot mobile phone middle frame high-adaptation positioning grabbing method
CN113771045A