High-speed workpiece interface positioning method based on scale invariant feature conversion
Through the method based on scale-invariant feature conversion and FPGA hardware acceleration, the calculation steps and resource consumption are optimized, and the problem of low positioning efficiency of workpiece interfaces in the prior art is solved, and efficient and accurate three-dimensional position detection of workpiece interfaces is achieved.
Patent Information
- Application Number
- CN202510677790.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-02
AI Technical Summary
The existing workpiece interface position detection technology consumes too much computing resources and time, which affects the detection efficiency. Especially in space-constrained scenarios, the system is more complex and it is difficult to achieve efficient and accurate workpiece interface positioning.
The method based on scale-invariant feature conversion is adopted, combined with FPGA hardware acceleration, and the three-dimensional position of the workpiece interface is calculated through binocular image acquisition, template matching, SIFT feature extraction and matching, and parallax method, and the calculation steps and hardware resource consumption are optimized to achieve real-time positioning.
It realizes efficient and accurate calculation of the three-dimensional position information of the workpiece interface on the FPGA platform, reducing the calculation amount and hardware resource consumption, and improving detection efficiency and accuracy.
Smart Images

Figure CN120580290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automated assembly, and in particular to a high-speed workpiece interface positioning method based on scale-invariant feature conversion. Background Art
[0002] In modern manufacturing, the efficiency and accuracy of automated workpiece assembly systems are directly related to production efficiency, product quality, and costs. Efficient object recognition and precise positioning are crucial for guiding assembly robotic arms. Therefore, developing efficient and accurate workpiece recognition and positioning methods is of great value. Current assembly interface position detection technologies primarily utilize computer vision combined with active vision technology using structured light projection, and the collaborative operation of cameras and multiple sensors. However, the high system complexity of these two methods limits their application in certain space-constrained scenarios.
[0003] Workpiece interface position detection technology is commonly used on robotic arms in assembly lines. This requires highly flexible position detection equipment capable of quickly and accurately obtaining workpiece interface position information. Binocular vision technology uses parallax to measure an object's depth, thereby obtaining highly accurate spatial position information. Compared to active vision and hybrid vision technologies, this technology offers simpler equipment. However, binocular vision target position measurement technology typically requires extensive computation and processing of the binocular images, resulting in a complex algorithm with numerous steps and significant computational resources and time consumption. This significantly impacts the efficiency of workpiece interface position detection.
[0004] To address the above problems, this paper proposes a high-speed workpiece interface positioning method based on scale-invariant feature transformation. Summary of the Invention
[0005] (1) Technical problems solved
[0006] In view of the deficiencies in the prior art, the present invention provides a high-speed workpiece interface positioning method based on scale-invariant feature transformation, which solves the problems raised in the above background technology.
[0007] (2) Technical solution
[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0009] A high-speed workpiece interface positioning method based on scale-invariant feature transformation includes the following steps:
[0010] S1: binocular image acquisition and distortion correction;
[0011] S2: Target recognition of the left image based on the template matching algorithm;
[0012] S3: Perform SIFT feature extraction and matching on the left and right images;
[0013] S4: Calculate the target's three-dimensional position information based on the parallax method;
[0014] S5: The above workpiece interface position detection method is hardware accelerated and optimized in FPGA to obtain real-time target position information.
[0015] Furthermore, in S1, the image acquisition system used in the high-speed workpiece interface positioning method based on scale-invariant feature transformation includes:
[0016] Binocular Camera Bracket: Used to fix two industrial cameras so that they have a fixed baseline distance.
[0017] Two industrial cameras: Two industrial cameras are used to collect binocular images of the workpiece;
[0018] FPGA development board: used to send control signals to the binocular camera and receive and process the binocular images captured by the binocular camera.
[0019] The camera is driven by FPGA to capture binocular images, converting RGB images into grayscale images, and quickly correcting lens distortion through a lookup table.
[0020] Furthermore, in S2, the left image of the corrected binocular image is first subjected to filtering and denoising. The filtering process adopts a guided filtering with linear smoothing characteristics, and the formula is as follows:
[0021]
[0022]
[0023] Where I is the original image; q is the output image; is the mean of I in the filter window; var(I) is the variance of the guided graph. Hardware-oriented optimization of the guided filter is performed by replacing the box filter with a Gaussian filter and quantizing the γ, ε, and ω parameters to constants, where ε / ω = 0.0001 and γ = 5000.
[0024] Denoising uses corrosion and expansion methods to eliminate noise. The formula is as follows:
[0025]
[0026] Where A is the target in the original input image; B is the eroded and expanded structural element, (B) z is the set of structural elements after translation z on the target image; A c is the complement of A; is the set of inverted structural elements B at the origin. After filtering and denoising the left image, the Canny edge detection algorithm is used to detect the edge of the image. The process is as follows:
[0027] (1) Use Gaussian filter to smooth the image and remove noise.
[0028] (2) Calculate the gradient intensity and direction of pixel points in the grayscale image.
[0029] (3) Eliminate spurious responses caused by edge detection.
[0030] (4) Detect real and potential edges. Suppress isolated weak edges and finally complete edge detection. After edge detection of the left image, the image edge is binarized and then SIFT-based feature point extraction and matching is performed with the matching template image. The matching template image is the workpiece interface image that has undergone the same filtering, denoising, and edge detection process as the left image. The RANSAC algorithm is used to eliminate mismatched points, and the upper, lower, left, and right boundaries of the feature points in the left image are used to generate a rectangular area, which is the area where the identified workpiece interface is located.
[0031] Furthermore, in S3, the hardware-optimized SIFT algorithm is used to extract and match feature points from the left and right images. The traditional SIFT algorithm's scale space structure of more than 8 groups is improved to a scale space structure of 1 group and 4 layers. Constructing the scale space requires convolving the input image with Gaussian kernels of different scale factors. The Gaussian kernel formula is as follows:
[0032]
[0033] The above formula can be further decomposed into:
[0034]
[0035] It can be seen from the above formula that a two-dimensional Gaussian function can be split into the product of two one-dimensional Gaussian functions, so a one-dimensional filter kernel can be time-division multiplexed into two filter kernels to save hardware resources. According to the normal distribution characteristics of the Gaussian filter kernel function, data far from the center point has little effect on the center point. Therefore, the present invention adopts a Gaussian kernel template with a fixed size of 7×7, and the scale parameters of each layer are σ1=1.6, σ2=1.226, σ3=1.545, and σ4=1.947. When constructing the feature descriptor, the area around the key point is divided into 2×2 sub-regions, each sub-region contains 3σ image elements, and the gradient information of each sub-region is evenly divided into 8 ranges according to the angle of 0~2π. The gradient information is statistically analyzed in each range to obtain a 32-dimensional feature descriptor. After that, the feature points of the left and right images are matched, and only the feature points within the rectangular area generated by step S2 are retained.
[0036] Furthermore, in S4, the three-dimensional position information of the object is calculated based on the parallax method, and the calculation model of binocular stereo vision is:
[0037]
[0038] The spatial position relationship between the two cameras can be expressed as:
[0039]
[0040] Where R is the transformation matrix between the two cameras; T is the translation vector between the two camera coordinate systems. Z can be calculated by the above formula. L The value of , thus obtaining the three-dimensional coordinates of the space point.
[0041] Furthermore, the entire process is implemented through FPGA, which mainly includes the following modules:
[0042] Clock module: used to generate and distribute periodic clock signals to ensure that all modules in the entire system are in a synchronized state in order to coordinate data transmission and processing;
[0043] Image acquisition module: used to send control signals to two industrial cameras and receive binocular images;
[0044] Distortion correction module: used to correct image distortion caused by lens distortion
[0045] Image cache module: used to temporarily store intermediate results or data that needs to be processed later to ensure that the data is not lost and can be accessed at the appropriate time;
[0046] Target detection module: used to identify the position information of the workpiece interface in the left image of the binocular image.
[0047] Stereo matching module: used to extract and match feature points in binocular images and calculate the target's three-dimensional position information based on the parallax method.
[0048] Data output module: used to output image data and workpiece interface position information to the host computer for further processing.
[0049] (3) Beneficial effects
[0050] Compared with the prior art, the present invention provides a high-speed workpiece interface positioning method based on scale-invariant feature transformation, which has the following beneficial effects:
[0051] The present invention improves the limitations of the traditional workpiece interface positioning method based on structured light projection combined with machine vision, and proposes a high-speed workpiece interface positioning method based on scale-invariant feature transformation, which can calculate the three-dimensional position information of the workpiece interface at high speed and accuracy.
[0052] Based on the implementation of the position detection algorithm on the FPGA platform, the SIFT algorithm is optimized for hardware. By optimizing the structure of the scale space, the amount of calculation is reduced, and the Gaussian filter kernel module for constructing the scale space is designed for time-sharing multiplexing to reduce the consumption of hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Schematic diagram of the flow of the workpiece interface position detection method used in an embodiment of the present invention;
[0054] Figure 2 The hardware structure of the time-division multiplexed Gaussian filter core proposed in the embodiment of the present invention;
[0055] Figure 3 This is an overall hardware block diagram of the workpiece interface position detection and extraction method in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0057] Example
[0058] See Figure 1-3 An embodiment of the present invention provides a high-speed workpiece interface positioning method based on scale-invariant feature transformation, the method specifically comprising steps S1-S5:
[0059] S1: binocular image acquisition and correction;
[0060] The image acquisition system used in the high-speed workpiece interface positioning method based on scale-invariant feature transformation includes:
[0061] Binocular Camera Bracket: Used to fix two industrial cameras so that they have a fixed baseline distance.
[0062] Two industrial cameras: Two industrial cameras are used to collect binocular images of the workpiece;
[0063] FPGA development board: used to send control signals to the binocular camera and receive and process the binocular images captured by the binocular camera.
[0064] The camera is driven by FPGA to capture binocular images, converting RGB images into grayscale images, and quickly correcting lens distortion through a lookup table.
[0065] S2: Target recognition is performed on the left image based on the template matching algorithm.
[0066] First, the left image of the corrected binocular image is filtered and denoised. The filtering process uses a guided filter with linear smoothing characteristics, and the formula is as follows:
[0067]
[0068] Where I is the original image; q is the output image; is the mean of I in the filter window; var(I) is the variance of the guided graph. Hardware-oriented optimization of the guided filter is performed by replacing the box filter with a Gaussian filter and quantizing the γ, ε, and ω parameters to constants, where ε / ω = 0.0001 and γ = 5000.
[0069] Denoising uses corrosion and expansion methods to eliminate noise. The formula is as follows:
[0070]
[0071] Where A is the target in the original input image; B is the eroded and expanded structural element, (B) z is the set of structural elements after translation z on the target image; A c is the complement of A; is the set of inverted structural elements B at the origin. After filtering and denoising the left image, the Canny edge detection algorithm is used to detect the edge of the image. The process is as follows:
[0072] (1) Use Gaussian filter to smooth the image and remove noise.
[0073] (2) Calculate the gradient intensity and direction of pixel points in the grayscale image.
[0074] (3) Eliminate spurious responses caused by edge detection.
[0075] (4) Detect real and potential edges. Suppress isolated weak edges and finally complete edge detection. After edge detection of the left image, the image edge is binarized and then SIFT-based feature point extraction and matching is performed with the matching template image. The matching template image is the workpiece interface image that has undergone the same filtering, denoising, and edge detection process as the left image. The RANSAC algorithm is used to eliminate mismatched points, and the upper, lower, left, and right boundaries of the feature points in the left image are used to generate a rectangular area, which is the area where the identified workpiece interface is located.
[0076] S3: Perform SIFT feature extraction and matching on the left and right images.
[0077] The hardware-optimized SIFT algorithm is used to extract and match feature points from the left and right images. The traditional SIFT algorithm's scale space structure of more than eight groups is improved to a single scale space structure with four layers. Constructing the scale space requires convolving the input image with Gaussian kernels of varying scale factors. The Gaussian kernel formula is as follows:
[0078]
[0079] The above formula can be further decomposed into:
[0080]
[0081] From the above formula, we can see that a two-dimensional Gaussian function can be split into the product of two one-dimensional Gaussian functions. Therefore, a one-dimensional filter kernel can be time-division multiplexed into two filter kernels to save hardware resources. Figure 2 This is the Gaussian filter kernel structure in the real-time example of the present invention. According to the normal distribution characteristics of the Gaussian filter kernel function, data far from the center point has little effect on the center point. Therefore, the present invention adopts a Gaussian kernel template with a fixed size of 7×7, and the scale parameters of each layer are σ1=1.6, σ2=1.226, σ3=1.545, and σ4=1.947. When constructing the feature descriptor, the area around the key point is divided into 2×2 sub-regions, each sub-region contains 3σ image elements, and the gradient information of each sub-region is evenly divided into 8 ranges according to the angle of 0~2π. The gradient information is statistically analyzed within each range to obtain a 32-dimensional feature descriptor. After that, the feature points of the left and right images are matched, and only the feature points within the rectangular area generated by step S2 are retained.
[0082] See also Figure 2
[0083] S4: Calculate the three-dimensional position information of the target based on the parallax method.
[0084] The three-dimensional position information of an object is calculated based on the parallax method. The calculation model of binocular stereo vision is:
[0085]
[0086] The spatial position relationship between the two cameras can be expressed as:
[0087]
[0088] Where R is the transformation matrix between the two cameras; T is the translation vector between the two camera coordinate systems. Z can be calculated by the above formula. L The value of , thus obtaining the three-dimensional coordinates of the space point.
[0089] S5: The workpiece interface position detection method is hardware accelerated and optimized in FPGA to obtain real-time target position information.
[0090] See also Figure 3 , shown is the overall hardware block diagram of the workpiece interface position detection method in an embodiment of the present invention.
[0091] The entire process is implemented through FPGA and mainly includes the following modules:
[0092] Clock module: used to generate and distribute periodic clock signals to ensure that all modules in the entire system are in a synchronized state in order to coordinate data transmission and processing;
[0093] Image acquisition module: used to send control signals to two industrial cameras and receive binocular images;
[0094] Distortion correction module: used to correct image distortion caused by lens distortion
[0095] Image cache module: used to temporarily store intermediate results or data that needs to be processed later to ensure that the data is not lost and can be accessed at the appropriate time;
[0096] Target detection module: used to identify the position information of the workpiece interface in the left image of the binocular image.
[0097] Stereo matching module: used to extract and match feature points in binocular images and calculate the target's three-dimensional position information based on the parallax method.
[0098] Data output module: used to output image data and workpiece interface position information to the host computer for further processing.
[0099] The present invention provides a high-speed workpiece interface location method based on scale-invariant feature transformation. To achieve this goal, this technical solution uses a template matching algorithm to identify the position of the workpiece interface in binocular images, and a hardware-optimized SIFT algorithm to extract and match feature points of the workpiece interface. The entire process is implemented through FPGA hardware acceleration. This method can efficiently, reliably, and accurately complete the workpiece interface position detection.
[0100] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A high-speed workpiece interface positioning method based on scale-invariant feature transformation, characterized in that: The following steps are included: S1: binocular image acquisition and distortion correction; S2: Target recognition of the left image based on the template matching algorithm; S3: Perform SIFT feature extraction and matching on the left and right images; S4: Calculate the target's three-dimensional position information based on the parallax method; S5: The above workpiece interface position detection method is hardware accelerated and optimized in FPGA to obtain real-time target position information.
2. The high-speed workpiece interface positioning method based on scale-invariant feature transformation according to claim 1 is characterized in that: In S1, the binocular image acquisition system includes: Binocular camera bracket: used to fix two industrial cameras so that they have a fixed baseline distance; Two industrial cameras: Two industrial cameras are used to collect binocular images of the workpiece; FPGA development board: used to send control signals to the binocular camera and receive and process the binocular images captured by the binocular camera; The camera is driven by FPGA to capture binocular images, converting RGB images into grayscale images, and quickly correcting lens distortion through a lookup table.
3. The high-speed workpiece interface positioning method based on scale-invariant feature transformation according to claim 1 is characterized in that: In S2, the left image of the corrected binocular image is first subjected to filtering and denoising. The filtering process adopts a guided filtering with linear smoothing features, and the formula is as follows: Where I is the original image; q is the output image; I is the value of I in the filter window; var(I) is the variance of the guided image. The guided filter is optimized for hardware, using Gaussian filtering instead of box filtering, and the parameters γ, ε, and ω are constantized, where ε / ω = 0.0001 and γ = 5000. Denoising uses corrosion and expansion methods to eliminate noise. The formula is as follows: Where A is the target in the original input image; B is the eroded and expanded structural element, (B) z is the set of structural elements after translation z on the target image; A c is the complement of A; is the set of inverted structural elements B at the origin. After filtering and denoising the left image, the Canny edge detection algorithm is used to detect the edge of the image. The process is as follows: (1) Use Gaussian filter to smooth the image and remove noise; (2) Calculate the gradient intensity and direction of the pixel points in the grayscale image; (3) Eliminate spurious responses caused by edge detection; (4) Detect real and potential edges, suppress isolated weak edges and finally complete edge detection; After edge detection of the left image, the image edges are binarized, and then SIFT-based feature point extraction and matching are performed with the matching template image. The matching template image is an image of the workpiece interface that has undergone the same filtering, denoising, and edge detection process as the left image. The RANSAC algorithm is used to eliminate false matching points, and a rectangular area is generated by the upper, lower, left, and right boundaries of the feature points in the left image. This area is the area where the identified workpiece interface is located.
4. The high-speed workpiece interface positioning method based on scale-invariant feature transformation according to claim 1 is characterized in that: In S3, the hardware-optimized SIFT algorithm is used to extract and match feature points from the left and right images. The traditional SIFT algorithm's scale space structure of more than 8 groups is improved to a scale space structure of 1 group and 4 layers. Constructing the scale space requires convolution of the input image with Gaussian kernels of different scale factors. The Gaussian kernel formula is as follows: The above formula can be further decomposed into: It can be seen from the above formula that a two-dimensional Gaussian function can be decomposed into the product of two one-dimensional Gaussian functions. Therefore, a one-dimensional filter kernel can be time-division multiplexed into two filter kernels to save hardware resources. According to the normal distribution characteristics of the Gaussian filter kernel function, data far from the center point has little effect on the center point. Therefore, a Gaussian kernel template with a fixed size of 7×7 is used, and the scale parameters of each layer are σ1=1.6, σ2=1.226, σ3=1.545, and σ4=1.
947. When constructing the feature descriptor, the area around the key point is divided into 2×2 sub-regions, each sub-region contains 3σ image elements, and the gradient information of each sub-region is evenly divided into 8 ranges according to the angle of 0~2π. The gradient information is statistically analyzed in each range to obtain a 32-dimensional feature descriptor. Then, the feature points of the left and right images are matched, and only the feature points within the rectangular area generated in step S2 are retained.
5. The high-speed workpiece interface positioning method based on scale-invariant feature transformation according to claim 1 is characterized in that: In S4, the three-dimensional position information of the object is calculated based on the parallax method, and the calculation model of binocular stereo vision is: The spatial position relationship between the two cameras can be expressed as: Where R is the transformation matrix between the two cameras; T is the translation vector between the two camera coordinate systems. Z can be calculated by the above formula. L , thereby obtaining the three-dimensional coordinates of the space point.
6. The high-speed workpiece interface positioning method based on scale-invariant feature transformation according to claim 1 is characterized in that: In S5, the entire process is implemented by FPGA, which mainly includes the following modules: Clock module: used to generate and distribute periodic clock signals to ensure that all modules in the entire system are in a synchronized state in order to coordinate data transmission and processing; Image acquisition module: used to send control signals to two industrial cameras and receive binocular images; Distortion correction module: used to correct image distortion caused by lens distortion; Image cache module: used to temporarily store intermediate results or data that needs to be processed later to ensure that the data is not lost and can be accessed at the appropriate time; Target detection module: used to identify the position information of the workpiece interface in the left image of the binocular image; Stereo matching module: used to extract and match feature points in binocular images and calculate the target's three-dimensional position information based on the parallax method; Data output module: used to output image data and workpiece interface position information to the host computer for further processing.