Local-posture-invariable template matching, assembling and positioning method, system and equipment
Through the template matching assembly positioning method with unchanged local attitude, the problems of insufficient accuracy, low efficiency and weak generalization ability in the prior art are solved, and high-precision, efficiency and robust robot assembly are achieved, suitable for complex and deformed workpiece environments.
Patent Information
- Application Number
- CN202510552266.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing robot vision-guided assembly technology has shortcomings in terms of high accuracy, efficiency and robustness, especially when dealing with complex or variable assembly tasks. The traditional global template matching method is sensitive to image noise, local occlusion and lighting changes, making it difficult to achieve high positioning accuracy, and has low efficiency, poor real-time performance, and weak generalization ability.
The template matching assembly positioning method with unchanged local posture is adopted, and the global coarse positioning is performed through the Sobel operator. The key points of the workpiece are extracted in combination with the scale-invariant feature transformation algorithm, and the local feature descriptor is generated. The position information of the workpiece is estimated by matching feature points, and the calibration information of the suction cup holder is combined to calculate the assembly positioning position at the end of the robot.
It improves the accuracy and efficiency of robot assembly, enhances the robustness and generalization capabilities of the system, can better adapt to complex and deformed workpiece environments, reduces manual intervention, and achieves high-precision assembly goals.
Smart Images

Figure CN120070454A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot vision-guided assembly, and particularly relates to a template matching assembly positioning method, system and device with invariant local pose. Background Art
[0002] Robot assembly is gradually playing an increasingly important role in high-precision operations such as product manufacturing, precision electronic product processing, and aerospace component assembly. With the development of industrial automation and intelligent manufacturing technologies, traditional rigid assembly methods are difficult to meet the requirements of flexible and customized production. The robot assembly technology based on vision guidance has become one of the key technologies in modern manufacturing due to its high precision, flexibility and self-adaptive ability. Traditional robot assembly heavily relies on predefined paths and mechanical fixtures, which limits the flexibility and adaptability when dealing with complex or variable assembly tasks. In contrast, vision-guided assembly can achieve real-time perception and adaptive control, significantly improving precision, efficiency and versatility. It provides an effective solution for automating the assembly process and improving the assembly quality. However, in the process of vision-guided automated assembly, errors may be introduced by sensors, robot joints, and end effector tools, and the accumulation of these errors still poses a challenge to the high-precision implementation of robot assembly.
[0003] The core challenge of robot vision-guided assembly technology lies in how to achieve high-precision, high-efficiency and robust workpiece positioning. Existing technologies mainly rely on global feature matching or contact detection methods, but still have the following significant defects: 1. Insufficient precision and error superposition Traditional global template matching methods (such as normalized cross-correlation matching, edge contour matching) are sensitive to image noise, local occlusion and illumination changes, and it is difficult to achieve high positioning precision. The superposition of the robot's own repeated positioning error and the absolute positioning error of the vision system (caused by camera calibration deviation, image distortion, etc.) further reduces the overall system precision.
[0004] 2. Low efficiency and poor real-time performance The global search algorithm needs to traverse the entire image, with high computational complexity. Taking an image with a resolution of 1280×1024 as an example, the full-image edge matching takes more than 200 ms, which is difficult to meet the real-time requirement (usually required to be ≤50 ms) of high-speed production lines. Contact detection (such as mechanical probe positioning) requires physical contact with the workpiece, and the single detection time increases by 30~50 ms, and it may scratch the surface of the precision workpiece.
[0005] 3. Weak generalization ability Existing methods have poor adaptability to changes in the shape, pose, and distance of workpieces. Methods based on single contour matching cannot handle workpiece rotation or scaling (e.g., the matching failure rate > 60% when the rotation exceeds ±30°); methods based on color segmentation fail under light fluctuations or background interference (e.g., a ±20% change in light intensity results in a 40% increase in the false detection rate). When switching between workpieces of multiple varieties, it is necessary to redesign the positioning algorithm, resulting in high transformation costs for flexible production lines.
[0006] Based on this, a template matching assembly positioning method, system, and device with local pose invariance are proposed. Summary of the Invention
[0007] In view of the above technical problems, the present invention provides a template matching assembly positioning method, system, and device with local pose invariance.
[0008] The technical solution adopted by the present invention to solve its technical problems is as follows: A template matching assembly positioning method with local pose invariance, the method comprising the following steps: S100: Collect image data of the workpiece to be assembled and the assembly environment; S200: Apply the Sobel operator to perform global rough positioning on the image data, detect edge information in the image through the gradient changes of horizontal and vertical edges, and determine the candidate region of the workpiece; S300: In the candidate region, use the scale-invariant feature transform algorithm to extract key points of the workpiece, and generate local feature descriptors based on the key points; S400: Match the local feature descriptors with the features of the pre-stored template, and estimate the pose information of the workpiece through the matched feature points; S500: Based on the constraint conditions of the fixed points of the suction cup for clamping the workpiece, combined with the transformation matrix from the coordinate system C of the pre-calibrated nozzle vacuum gripper to the workpiece coordinate system W, calculate the assembly pose of the robot end, and control the robot to complete the assembly operation of the workpiece according to the assembly pose.
[0009] Preferably, the Sobel operator in S200 includes a horizontal edge detection convolution kernel and a vertical edge detection convolution kernel, which are respectively used to detect the gradient changes in the horizontal and vertical directions of the image. S200 includes: S210: For each pixel point, by convolving the two convolution kernels in the Sobel operator with the input image, obtain the gradient images of the image in the horizontal and vertical directions and ; S220: According to the gradient images of the image in the horizontal and vertical directions and calculate the gradient magnitude M and the gradient direction , specifically: ; S230: Consider all points with gradient magnitudes less than a preset threshold T as the background, and all points with gradient magnitudes greater than the preset threshold T as edges. Extract the edges in the image, and represent the edge information with a binary image, where the edge part is white and the other parts are black.
[0010] Preferably, S300 includes: S310: Construct a Gaussian pyramid by continuously blurring and downsampling the image to generate a multi-scale image representation; S320: Detect scale-space extreme points on the images at different scales through the Gaussian pyramid, so as to find the key points with local scale invariance in the image; S330: In order to make the key points with scale invariance have rotational invariance, determine the main direction of each key point based on the gradient direction information of the local area of the image; S340: With the key point as the center, divide the local area into multiple small blocks. For each small block, calculate the direction and magnitude of the gradient, construct an orientation histogram, merge the gradient orientation histograms of all small blocks, generate a descriptor containing 128 elements, and finally normalize the 128-dimensional descriptor to make it have unit length.
[0011] Preferably, S310 is specifically: Apply a Gaussian filter to blur the image: ; where, is the image after Gaussian blur, represents the scale of the current layer, and respectively represent the coordinates of the input image; Downsample the blurred image by a certain ratio to generate the image of the next layer; Repeat the image blurring and downsampling process, and each time build a new layer on the basis of the image of the previous layer. Finally, the obtained Gaussian pyramid contains images of multiple scales; S320 includes: Generate a Difference of Gaussian (DoG) image by subtracting adjacent-scale Gaussian images. Specifically: ; where k is the scale increase factor, is the Difference of Gaussian image; Take the A pixel in the image is compared with the pixels of its adjacent 26 points. If the pixel value of the current pixel point is greater than or less than the pixel values of all 26 neighborhood points, it is a local maximum point or a local minimum point.
[0012] Preferably, S330 includes: S331: On the scale image where the key point is located, calculate the gradient magnitude and direction of the local area, specifically: ; ; Among them, and are the gradients of the image in the x and y directions respectively, θ is the gradient direction, and M is the gradient magnitude; S332: Calculate the weighted sum of the gradient magnitude and gradient direction within each direction interval of the key point, and select the direction with the largest peak in the histogram as the main direction of the key point.
[0013] Preferably, S400 includes: S410: Calculate the Euclidean distance between the local feature descriptor in the image to be matched and the feature descriptor of the pre-stored template image, specifically: ; Among them, and are the values of the i-th element in the local feature descriptor D1 and the feature descriptor D2 of the pre-stored template image, and n is the dimension in the descriptor; Select the descriptor pair with the smallest distance as the matching point; S420: Use the RANSAC algorithm to estimate the homography transformation matrix by randomly selecting a preset number of pairs of matching points, and obtain the estimated homography transformation matrix , and use the estimated homography transformation matrix to judge whether other matching points are consistent with the current transformation; repeat random sampling, and finally obtain a maximum inlier set, so as to effectively eliminate mismatched points; S430: Estimate the pose of the object by calculating the homography transformation matrix among the matching points in the maximum inlier set to minimize the error of the matching points.
[0014] Preferably, S420 includes: Use the RANSAC algorithm to randomly select 4 pairs of matching points from all the matching point sets to form a sampling set for estimating the homography transformation matrix, and obtain the estimated homography transformation matrix , specifically: ; Calculate the reprojection error of all matching points through the estimated homography transformation matrix : ; where, is the original image coordinate actually detected, is the target image coordinate that should be mapped theoretically according to the estimated homography transformation matrix; If the reprojection error is less than the set threshold, the matching point is an inlier, otherwise it is an outlier. Repeat the above process to finally determine the largest inlier set and eliminate the mismatched points; S430 includes: Calculate the homography transformation matrix using all the matching points in the largest inlier set , specifically: ; where, is the point in the original image of the j-th pair of matching points , is the transformed point of the j-th pair of matching points , and m is the total number of matching points; where, for the matching point pairs {( , )} and {( , )}, the homography transformation matrix satisfies the following relationship: ; is in the form of: ; where, the homography transformation matrix describes the transformation under different perspective views on the same plane, , , , control rotation, scaling, and shear transformation; , control translation transformation, , control perspective distortion, is the normalization factor; Decompose the homography transformation matrix to obtain the rotation matrix R and the translation vector T: ; Restore the pose of the object in 3D space through the decomposed rotation matrix R and translation vector T.
[0015] Preferably, S500 includes: S510: Define the pose of the vacuum gripper at the end of the robot in the base coordinate system B, where the pose of the vacuum gripper at the end includes position and orientation. S520: Calibrate the transformation matrix from the coordinate system C of the nozzle vacuum gripper to the workpiece coordinate system W , that is, transform the feature points from the gripper coordinate system C to the workpiece coordinate system W. S530: Determine the pose of the workpiece in the gripper coordinate system based on the position and orientation information of the vacuum gripper at the end of the robot, combined with the fixed constraint conditions between the nozzle vacuum gripper and the workpiece. The cooperation between the nozzle and the workpiece needs to meet the double coordinate system alignment conditions: ; Among them, is the fixed mating pose, represents the expected orientation of the gripper coordinate system relative to the workpiece coordinate system when the nozzle contacts the workpiece, is the expected coordinate of the center origin of the gripper in the workpiece coordinate system when the nozzle contacts the workpiece; S540: According to the transformation matrix from the nozzle vacuum gripper coordinate system C to the workpiece coordinate system W, combined with the pose of the workpiece in the base coordinate system, calculate the actual pose of the vacuum gripper at the end of the robot in the base coordinate system: ; Among them, is the pose of the gripper in the base coordinate system, is the pose of the workpiece in the base coordinate system, is the transformation matrix from the nozzle vacuum gripper coordinate system to the workpiece coordinate system; S550: According to the calculated pose information of the end effector, the robot adjusts its motion trajectory to ensure that the gripper is correctly fixed at the specified position of the workpiece, and performs correction in combination with the feedback of the machine vision system, so as to ensure the accurate assembly pose.
[0016] The template matching assembly positioning system with local pose invariance includes an image data acquisition module, a candidate region determination module, a local feature descriptor generation module, a workpiece pose information estimation module, and a workpiece assembly module; The image data acquisition module is used to acquire the image data of the workpiece to be assembled and the assembly environment; The candidate region determination module is used to perform global rough positioning on the image data by applying the Sobel operator, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate region of the workpiece; The local feature descriptor generation module is used to extract the key points of the workpiece in the candidate region by using the scale-invariant feature transform algorithm, and generate local feature descriptors based on the key points; The workpiece pose information estimation module is used to match the local feature descriptors with the features of the pre-stored template, and estimate the pose information of the workpiece through the matched feature points; The workpiece assembly module is used to solve the assembly pose of the robot end based on the constraint conditions of the suction cup clamping the fixed points of the workpiece, combined with the transformation matrix from the pre-calibrated coordinate system C of the suction nozzle vacuum gripper to the workpiece coordinate system W, and control the robot to complete the workpiece assembly operation according to the assembly pose.
[0017] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the template matching assembly positioning method with local pose invariance are implemented.
[0018] The above-mentioned template matching assembly positioning method, system and device with local pose invariance is an advanced method that combines precise image processing, feature matching and machine vision system with robot control, aiming to improve the accuracy and efficiency of the robot in the assembly process. Its advantages lie in high-precision positioning, strong robustness, multi-step processing optimization, dynamic adaptability and the close cooperation between the robot and the vision system. Technologies such as the RANSAC algorithm and suction cup gripper calibration are used to ensure the precise identification and assembly of the workpiece, reduce manual intervention, are widely applicable to high-precision assembly tasks, and have significant application potential. Description of the Drawings
[0019] Figure 1 It is a flowchart of a template matching assembly positioning method with local pose invariance in an embodiment of the present invention; Figure 2 It is a schematic diagram of obtaining robot image data provided in an embodiment of the present invention. Detailed Embodiments
[0020] In order to enable those skilled in the art of the present technology to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0021] In one embodiment, as Figure 1 shown, a template matching assembly positioning method with local pose invariance, the method includes the following steps: S100: Collect image data of the workpiece to be assembled and the assembly environment; S200: Apply the Sobel operator to perform global rough positioning on the image data, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate area of the workpiece; S300: In the candidate area, use the scale-invariant feature transform algorithm to extract the key points of the workpiece, and generate local feature descriptors based on the key points; S400: Match the local feature descriptors with the features of the pre-stored templates, and estimate the pose information of the workpiece through the matched feature points; S500: Based on the constraint conditions of the fixed points for the suction cup to hold the workpiece, combined with the pre-calibrated transformation matrix from the coordinate system C of the nozzle vacuum gripper to the workpiece coordinate system W, solve the assembly pose of the robot end, and control the robot to complete the workpiece assembly operation according to the assembly pose.
[0022] Compared with other assembly positioning methods, through local feature extraction and template matching, the present invention can accurately capture and locate the key points of the workpiece, greatly improving the positioning accuracy in terms of accuracy, robustness, and adaptability. Especially in a complex and deformed workpiece environment, the use of local pose invariant descriptors can reduce the errors caused by object pose changes or different perspectives, thereby effectively overcoming the limitations of traditional methods and achieving accurate robot assembly positioning.
[0023] Further, as Figure 2 shown in the figure, in the figure, B is the robot base coordinate system, C is the coordinate system of the nozzle vacuum gripper, W is the workpiece coordinate system, S is the camera. After the robot picks up the workpiece on the carrier plate, it moves to the lower camera photographing position to obtain an image. After determining the camera parameters using the calibration method, execute the above S100 - S500.
[0024] In one embodiment, the Sobel operator in S200 includes a horizontal edge detection convolution kernel and a vertical edge detection convolution kernel, which are respectively used to detect the gradient changes in the horizontal and vertical directions of the image. S200 includes: S210: For each pixel point, by performing convolution operations on the input image (grayscale image) with the two convolution kernels in the Sobel operator, obtain the gradient images of the image in the horizontal and vertical directions and ; S220: According to the gradient images of the image in the horizontal and vertical directions and calculate the gradient magnitude M (edge intensity) and the gradient direction , specifically: ; S230: In order to better highlight the edges in the image, usually perform thresholding on the gradient magnitude. Consider all points with gradient magnitude less than the preset threshold T as the background (i.e., not belonging to the edge), and consider all points with gradient magnitude greater than the preset threshold T as the edge, extract the edges in the image, and represent the edge information with a binary image, where the edge part is white and other parts are black.
[0025] In one embodiment, S300 includes: S310: Construct a Gaussian pyramid by continuously blurring and downsampling the image to generate a multi-scale image representation; S320: Detect scale-space extreme points on the images at different scales of the Gaussian pyramid to find the key points with local scale invariance in the image; S330: To make the key points with scale invariance have rotational invariance, determine the main direction of each key point based on the gradient direction information of the local region of the image; S340: Centered on the key points, divide the local region into multiple small blocks. For each small block, calculate the direction and magnitude of the gradient, construct an orientation histogram, merge the gradient orientation histograms of all small blocks, generate a descriptor containing 128 elements, and finally normalize the 128-dimensional descriptor to make it have unit length.
[0026] In one embodiment, S310 specifically is: Apply a Gaussian filter to blur the image: ; Wherein, is the image after Gaussian blur, represents the scale of the current layer, and respectively represent the coordinates of the input image; Downsample the blurred image by a certain ratio to generate the image of the next layer; Repeat the image blurring and downsampling process, and each time build a new layer on the basis of the image of the previous layer. Finally, the obtained Gaussian pyramid contains images of multiple scales; S320 includes: Generate a difference-of-Gaussians (DoG) image by subtracting adjacent-scale Gaussian images. Specifically: ; Where k is the scale increase factor, is the difference-of-Gaussians image; Take a pixel in the image of the current scale and compare it with the pixels of its adjacent 26 points. If the pixel value of the current pixel point is greater than or less than the pixel values of all 26 neighborhood points, it is a local maximum point or a local minimum point.
[0027] In one embodiment, S330 includes: S331: On the scale image where the key point is located, calculate the gradient magnitude and direction of the local region. Specifically: ; ; Wherein, and are the gradients of the image in the x and y directions respectively, θ is the gradient direction, and M is the gradient magnitude; S332: Calculate the weighted sum of the gradient magnitude and gradient direction within each direction interval of the key point, and select the direction with the largest peak in the histogram as the main direction of the key point.
[0028] In one embodiment, S400 includes: S410: Calculate the Euclidean distance between the local feature descriptor in the image to be matched and the feature descriptor of the pre-stored template image, specifically: ; where, and are the values of the i-th element in the local feature descriptor D1 and the feature descriptor D2 of the pre-stored template image, and n is the dimension in the descriptor; Select the pair of descriptors with the smallest distance as the matching points; generally speaking, the smaller the Euclidean distance of the descriptors, the more similar these two feature points are; S420: Use the RANSAC algorithm to estimate the homography transformation matrix by randomly selecting a preset number of pairs of matching points, and obtain the estimated homography transformation matrix , and use the estimated homography transformation matrix to determine whether other matching points are consistent with the current transformation; repeat random sampling, and finally obtain a maximum inlier set, thereby effectively removing false matching points; S430: Estimate the pose of the object by calculating the homography transformation matrix among the matching points in the maximum inlier set to minimize the error of the matching points.
[0029] In one embodiment, S420 includes: Use the RANSAC algorithm to randomly select 4 pairs of matching points from all matching point sets to form a sampling set for estimating the homography transformation matrix, and obtain the estimated homography transformation matrix , specifically: ; Calculate the reprojection error of all matching points through the estimated homography transformation matrix : ; where, is the original image coordinate actually detected, is the target image coordinate that should be mapped theoretically according to the estimated homography transformation matrix; If the reprojection error is less than the set threshold, the matching point is an inlier; otherwise, it is an outlier. Repeat the above process to finally determine the largest inlier set and eliminate the mismatched points; S430 includes: Calculate the homography transformation matrix using all the matching points in the largest inlier set , specifically: ; Among them, is the point in the original image of the j-th pair of matching points , is the transformed point of the j-th pair of matching points , and m is the total number of matching points; Among them, for the matching point pairs {( , )} and {( , )}, the homography transformation matrix satisfies the following relationship: ; The form of ; Among them, the homography transformation matrix describes the transformation under different perspective views on the same plane, , , , control rotation, scaling, and shear transformations; , control translation transformation, , control perspective distortion, is the normalization factor, usually set to 1; Decompose the homography transformation matrix to obtain the rotation matrix R and the translation vector T: ; Restore the pose of the object in the three-dimensional space through the rotation matrix R and the translation vector T obtained by decomposition.
[0030] In one embodiment, S500 includes: S510: Define the pose of the end effector (suction nozzle vacuum gripper) of the robot in the base coordinate system. Among them, the pose of the end effector includes the position (x, y, z coordinates) and the attitude (rotation angle or rotation matrix); usually, the robot control system will use the homogeneous transformation matrix to represent this pose; S520: Calibrate the transformation matrix from the coordinate system C of the suction nozzle vacuum gripper to the workpiece coordinate system W; further, the purpose of calibrating the transformation matrix is to correlate the pose of the suction cup gripper with the pose of the workpiece. The transformation matrix is usually obtained through the calibration process, which describes the position and attitude differences between the gripper coordinate system and the workpiece coordinate system. Through calibration, a homogeneous transformation matrix from the gripper coordinate system to the workpiece coordinate system can be obtained. ; S530: Determine the position of the gripper in the workpiece coordinate system based on the position and attitude information of the robot end effector and the fixed constraint conditions between the suction nozzle vacuum gripper and the workpiece; further, the relative position and attitude of the suction nozzle on the workpiece are determined, the attitude of the robot when the suction nozzle picks up the workpiece to be assembled is determined, and the suction nozzle vacuum gripper is at the end of the robot. Therefore, the relative position between the workpiece and the suction nozzle is used as a fixed constraint condition. The cooperation between the suction nozzle and the workpiece needs to meet the double coordinate system alignment conditions: ; where, is the fixed mating pose, represents the expected attitude of the gripper coordinate system relative to the workpiece coordinate system when the suction nozzle contacts the workpiece, is the expected coordinate of the center origin of the gripper in the workpiece coordinate system when the suction nozzle contacts the workpiece; S540: According to the transformation matrix from the suction nozzle vacuum gripper coordinate system C to the workpiece coordinate system W, combined with the pose of the robot end effector in the base coordinate system, solve the actual pose of the robot end effector in the workpiece coordinate system: ; where, is the pose of the gripper in the base coordinate system, is the pose of the workpiece in the base coordinate system, is the transformation matrix from the suction nozzle vacuum gripper coordinate system to the workpiece coordinate system; S550: According to the calculated pose information of the end effector, the robot adjusts its motion trajectory to ensure that the gripper is correctly fixed at the specified position of the workpiece, and performs correction in combination with the feedback of the machine vision system to ensure an accurate assembly pose.
[0031] In a detailed embodiment, the present invention is tested on a SCARA robot. Image data of the workpiece to be assembled and its environment is collected. Global rough positioning is performed on the image data. The gradient information of the image is detected by the Sobel operator, and the image is thresholded based on the gradient magnitude and direction to extract the edge information in the image. Key points of the image are extracted through the Gaussian pyramid. Combining the extreme points in the scale space, key local feature points in the image can be detected. The main direction of each key point is calculated through the weighted histogram of the gradient direction. The extracted local pose-invariant descriptor is matched with the features of the pre-stored template, and the Euclidean distance and the RANSAC algorithm are used to screen the optimal matching points. The RANSAC algorithm randomly selects matching points to estimate the homography transformation matrix, thereby eliminating the mismatched points. After obtaining the pose of the workpiece, combined with the calibration information of the suction cup gripper, the pose of the gripper is docked with the coordinate system of the workpiece. Combining the feedback of the robot control system and the machine vision system, it is ensured that the robot can perform precise assembly operations according to the calculated assembly pose, reducing human intervention and achieving the high-precision assembly goal.
[0032] Table 1 Positioning offset
[0033] The above-mentioned template matching positioning method based on local pose invariance includes: collecting image data of the workpiece to be assembled and the assembly environment; implementing a high-precision assembly of the robot under visual guidance by combining the assembly strategy of global rough positioning and fine positioning of local pose invariance. A local feature extraction algorithm is used to extract key points in the image. Assuming that the suction cup can always suck a fixed point on the workpiece each time, the transformation from the gripper to the workpiece coordinate system is always a fixed value. By calibrating the transformation matrix from the workpiece to the end coordinates, the assembly coordinate pose of the robot can be indirectly determined; the workpiece position is located by matching and processing the information using local pose invariance. Through the matched feature points, pose information such as the relative position, angle, and rotation of the target object is estimated. The present invention combines visual information with the local reference pose of the robot to provide a new method for calculating the assembly pose, solving the problem of inaccurate repeat positioning accuracy of the robot body and absolute positioning accuracy of the system; the above technical solution overcomes the technical problems of low efficiency, poor real-time performance, and contact inspection in the existing planar positioning technology, effectively improving the detection efficiency and detection accuracy, and the method of the present invention can also solve the problem of poor generalization ability for workpieces with different shapes, poses, and distances.
[0034] A template matching assembly positioning system with local pose invariance includes an image data acquisition module, a candidate region determination module, a local feature descriptor generation module, a workpiece pose information estimation module, and a workpiece assembly module; The image data acquisition module is used to collect image data of the workpiece to be assembled and the assembly environment; A candidate region determination module is used to perform global rough positioning on image data by applying the Sobel operator, detect edge information in the image through the gradient changes of horizontal and vertical edges, and determine the candidate region of the workpiece; A local feature descriptor generation module is used to extract key points of the workpiece within the candidate region by using the scale-invariant feature transform algorithm and generate local feature descriptors based on the key points; A workpiece pose information estimation module is used to match the local feature descriptors with the features of a pre-stored template and estimate the pose information of the workpiece through the matched feature points; A workpiece assembly module is used to solve the assembly pose of the robot end based on the constraint conditions of the fixed points for the suction cup to hold the workpiece, in combination with the transformation matrix from the pre-calibrated coordinate system C of the suction nozzle vacuum gripper to the workpiece coordinate system W, and control the robot to complete the workpiece assembly operation according to the assembly pose.
[0035] For the specific limitations of a template matching assembly positioning system with local pose invariance, reference can be made to the limitations of a template matching assembly positioning method with local pose invariance in the above text, which will not be elaborated here. Each module in the above template matching assembly positioning system with local pose invariance can be implemented in whole or in part through software, hardware, and their combinations. The above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0036] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a template matching assembly positioning method with local pose invariance are implemented.
[0037] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0038] The above has introduced in detail a local pose-invariant template matching assembly positioning method, system, and device provided by the present invention. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A template matching assembly positioning method with local posture invariance, characterized in that: The method comprises the following steps: S100: Collecting image data of the workpiece to be assembled and the assembly environment; S200: Apply the Sobel operator to perform global rough positioning on the image data, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate area of the workpiece; S300: In the candidate area, a scale-invariant feature transformation algorithm is used to extract key points of the workpiece, and a local feature descriptor is generated based on the key points; S400: matching the local feature descriptor with the features of the pre-stored template, and estimating the pose information of the workpiece through the matched feature points; S500: Based on the constraint conditions of the fixed point of the suction cup clamping the workpiece, combined with the pre-calibrated transformation matrix from the nozzle vacuum gripper coordinate system C to the workpiece coordinate system W, the assembly posture of the robot end is solved, and the robot is controlled according to the assembly posture to complete the assembly operation of the workpiece.
2. The method according to claim 1, characterized in that The Sobel operator in S200 includes a horizontal edge detection convolution kernel and a vertical edge detection convolution kernel, which are used to detect the gradient changes in the horizontal and vertical directions of the image respectively. S200 includes: S210: For each pixel, the two convolution kernels in the Sobel operator are convolved with the input image to obtain the gradient image of the image in the horizontal and vertical directions. and ; S220: Based on the gradient image of the image in the horizontal direction and the vertical direction and Calculate the gradient magnitude M and gradient direction , specifically: ; S230: All points with gradient amplitudes less than a preset threshold value T are regarded as background, and all points with gradient amplitudes greater than the preset threshold value T are regarded as edges, and the edges in the image are extracted. The edge information is represented by a binary image, in which the edge part is white and the other parts are black.
3. The method according to claim 2, characterized in that S300 includes: S310: constructing a Gaussian pyramid by continuously blurring and downsampling the image to generate a multi-scale image representation; S320: Detecting scale space extreme points on images of different scales through Gaussian pyramids, thereby finding key points with local scale invariance in the image; S330: In order to make the key points with scale invariance have rotation invariance, determine the main direction of each key point based on the gradient direction information of the local area of the image; S340: With the key point as the center, the local area is divided into multiple small blocks. For each small block, the direction and amplitude of the gradient are calculated, a direction histogram is constructed, the gradient direction histograms of all small blocks are merged to generate a descriptor containing 128 elements, and finally the 128-dimensional descriptor is standardized to have unit length.
4. The method according to claim 3, characterized in that S310 specifically includes: applying a Gaussian filter to perform Gaussian blur on the image: ; in, is the image after Gaussian blur, Represents the scale of the current layer, and Respectively represent the coordinates of the input image; Downsample the blurred image at a certain ratio to generate the image of the next layer; Repeat the image blurring and downsampling process, each time building a new layer based on the previous layer of image, and finally obtain a Gaussian pyramid, which contains images of multiple scales; S320 includes: The difference Gaussian image is generated by subtracting Gaussian images of adjacent scales, specifically: ; Among them, k is the scale increase coefficient, is the difference Gaussian image; Take the current scale A pixel in the image is compared with the pixels of its 26 neighboring points. If the pixel value of the current pixel is greater or less than the pixel values of all 26 neighboring points, it is a local maximum point or a local minimum point.
5. The method according to claim 4, characterized in that S330 includes: S331: On the scale image where the key point is located, the gradient amplitude and direction of the local area are calculated, specifically: ; ; in, and are the gradients of the image in the x and y directions, θ is the gradient direction, and M is the gradient magnitude; S332: Calculate the weighted sum of the gradient amplitude and gradient direction in each direction interval of the key point, and select the direction with the largest peak value in the histogram as the main direction of the key point.
6. The method according to claim 5, characterized in that S400 includes: S410: Calculate the Euclidean distance between the local feature descriptor in the image to be matched and the feature descriptor of the template image stored in advance, specifically: ; in, and is the value of the i-th element in the local feature descriptor D1 and the feature descriptor D2 of the pre-stored template image, and n is the dimension in the descriptor; Select the descriptor pair with the smallest distance as the matching point; S420: Using the RANSAC algorithm, a preset number of matching points are randomly selected to estimate the homography transformation matrix, thereby obtaining an estimated homography transformation matrix , and use the estimated homography transformation matrix to determine whether other matching points are consistent with the current transformation; repeat random sampling to eventually obtain a maximum internal point set, thereby effectively eliminating mismatched points; S430: by calculating the homography transformation matrix between the matching points in the maximum inner point set The pose of the object is estimated by minimizing the error of matching points.
7. The method according to claim 6, characterized in that S420 includes: Using the RANSAC algorithm, randomly select 4 pairs of matching points from all matching point sets Composed sampling set, used to estimate the homography transformation matrix, to obtain the estimated homography transformation matrix , specifically: ; Calculate the reprojection error of all matching points using the estimated homography transformation matrix : ; in, is the original image coordinate actually detected, is the target image coordinate that should be mapped to according to the estimated homography transformation matrix in theory; If the reprojection error is less than the set threshold, the matching point is an internal point, otherwise it is an external point. The above process is repeated to finally determine the maximum internal point set and eliminate the wrong matching points. S430 includes: Calculate the homography transformation matrix using all matching points in the maximum internal point set , specifically: ; in, is the point of the original image in the jth pair of matching points , is the transformed point in the jth pair of matching points , m is the total number of matching points; Among them, for the matching point pair {( , )} and{( , )}, the homography transformation matrix satisfies the following relationship: ; The form is: ; Among them, the homography transformation matrix It describes the transformation under different perspectives on the same plane. , , , Control rotation, scaling, and shear transformations; , Controls the translation transformation, , Control perspective distortion, is the normalization factor; Decomposing the homography transformation matrix Get the rotation matrix R and translation vector T: ; The position and posture of the object in three-dimensional space are restored by decomposing the rotation matrix R and translation vector T.
8. The method according to claim 7, characterized in that S500 includes: S510: defining the posture of the robot end nozzle vacuum gripper in the base coordinate system B, wherein the posture of the end nozzle vacuum gripper includes position and posture; S520: Calibrate the transformation matrix from the nozzle vacuum gripper coordinate system C to the workpiece coordinate system W , that is, transform the feature points from the gripper coordinate system C to the workpiece coordinate system W; S530: Determine the position and posture of the workpiece in the gripper coordinate system through the position and posture information of the vacuum gripper of the robot end nozzle, combined with the fixed constraint conditions between the vacuum gripper of the nozzle and the workpiece. The coordination between the nozzle and the workpiece must meet the dual coordinate system alignment conditions: ; in, To fix the matching posture, It indicates the expected posture of the gripper coordinate system relative to the workpiece coordinate system when the nozzle contacts the workpiece. is the expected coordinate of the center origin of the gripper in the workpiece coordinate system when the nozzle contacts the workpiece; S540: Transformation matrix from nozzle vacuum gripper coordinate system C to workpiece coordinate system W , combined with the pose of the workpiece in the base coordinate system, calculate the actual pose of the robot end nozzle vacuum gripper in the base coordinate system: ; in, is the position of the gripper in the base coordinate system, is the position and posture of the workpiece in the base coordinate system, is the transformation matrix from the nozzle vacuum gripper coordinate system to the workpiece coordinate system; S550: Based on the calculated position and posture information of the end effector, the robot adjusts its motion trajectory to ensure that the gripper is correctly fixed at the specified position of the workpiece, and performs corrections in combination with the feedback from the machine vision system to ensure accurate assembly posture.
9. A template matching assembly positioning system with local posture invariance, characterized in that: It includes an image data acquisition module, a candidate region determination module, a local feature descriptor generation module, a workpiece posture information estimation module and a workpiece assembly module; An image data acquisition module is used to acquire image data of the workpiece to be assembled and the assembly environment; The candidate region determination module is used to apply the Sobel operator to perform global rough positioning on the image data, detect the edge information in the image through the gradient changes of the horizontal and vertical edges, and determine the candidate region of the workpiece; A local feature descriptor generation module is used to extract key points of the workpiece in the candidate area using a scale-invariant feature transformation algorithm and generate local feature descriptors based on the key points; The workpiece pose information estimation module is used to match the local feature descriptor with the features of the pre-stored template and estimate the pose information of the workpiece through the matched feature points; The workpiece assembly module is used to solve the assembly posture of the robot end based on the constraint conditions of the suction cup clamping the workpiece fixed point, combined with the pre-calibrated transformation matrix from the nozzle vacuum gripper coordinate system C to the workpiece coordinate system W, and control the robot to complete the assembly operation of the workpiece according to the assembly posture.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
A flexible assembly system and method based on multi-mode information description
CN109543823A
Visual guidance right-angle robot mobile phone middle frame high-adaptation positioning grabbing method
CN113771045A
Tube plate workpiece welding method and system
CN117484047A
Method for Registering Points and Planes of 3D Data in Multiple Coordinate Systems
US20140003705A1
Double-robot collaborative assembly system and method for large wing skeleton
WO2024199008A1