Target positioning method and system based on machine vision

By selecting machine vision key hardware, preprocessing, establishing and using internal and external parameter matrices, and camera calibration methods in machine vision target positioning technology, the problems of insufficient feature extraction stability, image acquisition hardware adaptability and coordinate system conversion accuracy in the prior art are solved, and higher target positioning accuracy and stability are achieved.

CN120088333AActive Publication Date: 2025-06-03SINOHYDRO JIAJIANG HYDRAULIC MACHINERY +1

Patent Information

Application Number
CN202510007326.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-03
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing machine vision target positioning technology has insufficient problems in feature extraction stability, image acquisition hardware adaptability, and coordinate system conversion accuracy.

Method used

Target positioning calibration is performed by performing machine vision key hardware selection, image preprocessing, establishing coordinate systems and generating internal and external parameter matrices, and using camera calibration methods. Specific steps include using the SIFT algorithm to detect key points and feature points of the image, establishing a pixel coordinate system and image coordinate system, generating internal and external parameter matrices and performing camera calibration to calibrate the camera's geometric distortion.

Benefits of technology

It improves the stability of image acquisition and the accuracy of feature extraction, enhances the accuracy of coordinate system conversion, reduces positioning errors, and improves the overall accuracy and stability of target positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088333A_ABST
    Figure CN120088333A_ABST
Patent Text Reader

Abstract

The invention discloses a target positioning method and system based on machine vision, and relates to the technical field of machine vision and target positioning, and the method comprises the steps: carrying out the model selection of machine vision key hardware, and collecting image information; carrying out image preprocessing, and automatically detecting key points and feature points of the image by using an SIFT algorithm; establishing a coordinate system to generate an internal reference matrix, establishing a camera coordinate system to generate an external reference matrix, and performing coordinate transformation by combining the internal reference matrix and the external reference matrix; through a camera calibration method, target positioning calibration is carried out. The SIFT algorithm adopted by the method has very strong robustness for scale change, rotation, illumination change and a certain degree of view angle change of the image, and is suitable for a complex environment in an industrial scene. In the calibration process, a chessboard calibration plate and a least square method are utilized to improve the calculation precision of internal and external parameters, the calibration error is effectively reduced, and the positioning precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of machine vision and target positioning, and specifically provides a target positioning method and system based on machine vision. Background Art

[0002] With the rapid development of industrial automation, intelligent manufacturing, and artificial intelligence, machine vision technology has become a key component in industrial applications and is widely used in fields such as robot navigation, quality inspection, and target recognition. Machine vision endows computers with the "vision" ability by collecting and processing image data, enabling them to identify and locate target objects. In recent years, the improvement of hardware performance and the continuous innovation of image processing algorithms have provided strong support for the popularization and development of machine vision technology. However, in practical applications of high-precision target positioning, there are still many challenges in the stability of image acquisition, the accuracy of feature extraction, and the conversion accuracy of multi-dimensional space coordinates. Different application scenarios have different requirements for target positioning. Generally, it is necessary to optimize the hardware configuration of image acquisition and select appropriate algorithms to extract key features in the image, so as to achieve stable and accurate target positioning.

[0003] Existing machine vision target positioning technologies mainly rely on a variety of image processing algorithms and camera calibration technologies to achieve the conversion from two-dimensional image data to three-dimensional space. Common feature extraction algorithms include SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), etc., which show high robustness in complex scenarios such as rotation and scale changes. However, these methods still have certain deficiencies in practical applications. First, traditional hardware configuration schemes lack flexibility and are difficult to adapt to complex and changeable industrial environments, resulting in the inability to stably guarantee the quality of image acquisition. Second, although feature extraction algorithms have strong robustness, their stability and accuracy may be limited in the case of significant noise interference and light changes. In addition, existing methods for establishing coordinate systems and generating internal and external parameter matrices often lead to a reduction in positioning accuracy due to the instability of camera calibration. Accurate and rapid camera calibration is crucial for improving target positioning accuracy, but existing technologies have certain limitations in calibration methods and parameter optimization and have not effectively solved the calibration error problem in dynamic application scenarios. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed.

[0005] Therefore, the technical problem to be solved by the present invention is: the deficiencies of existing machine vision target positioning technologies in terms of feature extraction stability, adaptability of image acquisition hardware, and coordinate system conversion accuracy.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A target positioning method based on machine vision, including

[0007] Select key hardware for machine vision and collect image information;

[0008] Perform image preprocessing, and use the SIFT algorithm to automatically detect key points and feature points of the image;

[0009] Establish a coordinate system to generate an internal parameter matrix, establish a camera coordinate system to generate an external parameter matrix, and perform coordinate transformation by combining the internal and external parameter matrices;

[0010] Perform target positioning and calibration by means of camera calibration.

[0011] As a preferred solution of the target positioning method based on machine vision according to the present invention, wherein: the selection of key hardware for machine vision includes selecting a light source and a lighting method, and determining the resolution of the camera and the lens;

[0012] For lens selection, the main consideration factor is the focal length; assume that the size of the field of view to be photographed is FOV, and the horizontal and vertical directions are FOV H and FOV V , the horizontal and vertical sizes of the target surface of the camera chip are S H and S V , the detection distance is d w , then the horizontal component calculation formula and the vertical component calculation formula of the focal length f are expressed as:

[0013]

[0014] wherein, f H represents the horizontal focal length size, and f V represents the vertical focal length size.

[0015] As a preferred solution of the target positioning method based on machine vision according to the present invention, wherein: the collection of image information includes using two cameras to collect images; the first camera collects the end face images of the pull rod round hole surface and the rear suspension shaft of the pull rod from the front, and detects the alignment of the pull rod holes and the alignment of the pull rod and the suspension shaft; the second camera collects the side images of the pull rod and the suspension shaft, and detects the alignment of the pull rod holes and the penetration degree of the suspension shaft pin.

[0016] As a preferred solution of the target positioning method based on machine vision according to the present invention, wherein: the image preprocessing includes gray processing based on color features, image filtering, and binary processing of the image to achieve rough ROI positioning;

[0017] Select a suitable gray transformation method to transform the three-channel color RGB image into a single-channel gray image with the same value range; set the gray value at the pixel point P(x, y) after gray transformation to Gray(x, y);

[0018] Average grayscale processing method. The average value obtained from the RGB three-channel component values is used as the grayscale value of the corresponding pixel after grayscale transformation. The formula is expressed as:

[0019]

[0020] Using the weighted average grayscale processing method, the average grayscale processing method is generalized. After assigning appropriate weights to the RGB three-channel component values respectively, the grayscale value of the corresponding pixel for grayscale transformation is calculated. The formula is expressed as:

[0021] Gray(x,y) = r 1 R(x,y) + r 2 G(x,y) + r 3 B(x,y)

[0022] Where, R(x,y) represents the pixel at the R channel, G(x,y) represents the pixel at the G channel, and B(x,y represents the pixel at the B channel; r 1 、r 2 、r 3 represent the initial weights assigned to each channel component;

[0023] The ROI rough positioning process is based on the grayscale image after the smoothing process in the previous section, and inverse binary threshold processing is performed. During the processing, the threshold thresh is selected, and the formula is expressed as:

[0024]

[0025] Where, src(x,y) represents the pixel value at the coordinate (x,y) in the input image; dst(x,y) represents the pixel value of the output image at the coordinate (x,y); maxval represents the maximum value to be assigned to those pixels that meet the conditions;

[0026] Obtain the center point of the largest connected component. From the binary image, it can be seen that the position of the largest connected component is the position where the target to be detected is located, and its center point is also approximately the center point of the area. After obtaining the center point position, ROI rough positioning can be performed according to the round hole size information;

[0027] Obtain the ROI rough positioning area. Based on the approximate position of the target center point obtained above, according to the size information, a rectangular frame obtained by expanding several pixel points from the center point to the surrounding is the rough positioning area containing the target to be detected;

[0028] Perform ROI rough positioning area cropping. Only the rough positioning area needs to be processed in the subsequent processing.

[0029] As a preferred solution of the object positioning method based on machine vision according to the present invention, wherein: the automatic detection of key points and feature points of the image using the SIFT algorithm includes performing scale-space extreme value detection, searching for image positions at all scales, and identifying potential scale- and rotation-invariant interest points through Gaussian differential functions;

[0030] Performing precise positioning of key points, at each candidate position, determining the position and scale through a finely fitted model, and the selection of key points is based on their stability;

[0031] Determining the main direction of key points, based on the local gradient direction of the image, assigning one or more directions to each key point position, and all subsequent operations on the image data are transformed relative to the direction, scale, and position of the key points, so as to provide invariance to these transformations;

[0032] Generating sift feature vectors, measuring the local gradient of the image at a selected scale within the neighborhood around each key point;

[0033] When performing scale-space extreme value detection, an additional dynamic feedback adjustment mechanism is added to monitor the number and distribution of detected key points in real time; if too few or too many key points are detected at a specific scale, the scale parameter of Gaussian blur is automatically adjusted.

[0034] As a preferred solution of the object positioning method based on machine vision according to the present invention, wherein: the establishment of a coordinate system to generate an internal parameter matrix includes that the pixel coordinate system is a coordinate system of the image with respect to the determined target point, with the basic unit being pixels, and it is established on the photosensitive element of the camera; the image coordinate system is a coordinate system with physical length as the unit, and it is also established on the photosensitive element of the camera;

[0035] Converting the image coordinates (u, v) to normalized camera coordinates (x n , y n ), and the formula is expressed as:

[0036]

[0037] wherein, f x and f y represent the pixel units of the camera focal length; c x and c y represent the pixel coordinates of the principal point of the image;

[0038] Converting the normalized camera coordinates to the camera coordinate system, mapping the normalized camera coordinates into the camera coordinate system, assuming that the depth Z c of the object is known, then the formula is expressed as:

[0039] X c = xn ·Z c , Y c = y n ·Z c , Z c = Z c

[0040] The camera coordinate system is a reference system that describes the orientation of the target based on the camera as a reference object and uses physical length as the unit. The X-axis direction and the Y-axis direction are defined to be the same as the x-axis and the y-axis respectively. The Z-axis is perpendicular to the camera's photosensitive element, and the direction away from the photosensitive element is taken as positive. The center of the lens is the origin of this coordinate system;

[0041] Integrating the entire process, the image coordinates (u, v) are converted into world coordinates X w , Y w , Z w ) by the following complete formula:

[0042]

[0043] Among them, K represents the internal parameter matrix

[0044]

[0045] Among them, s represents the scale factor, which is calculated according to the actual depth information.

[0046] As a preferred solution of the target positioning method based on machine vision described in the present invention, wherein: the calibration by combining the internal parameter and external parameter matrices includes performing camera calibration, selecting a suitable checkerboard calibration board for printing, and determining the physical dimensions of the relevant points on the calibration board to ensure the accuracy of the calibration board;

[0047] Fix the calibration board on a flat panel, and within the camera's field of view, change the position and angle of the calibration board and take multiple photos correspondingly;

[0048] Use an algorithm to detect the corner feature points of the checkerboard calibration board in the image;

[0049] Solve and optimize the distortion coefficients. After initially solving the internal and external parameters of the camera, consider the image distortion factor; establish a radial distortion and tangential distortion model:

[0050] x' = x + (x - c x )(k 1 r 2 + k 2 r 4 ) + 2p 1 xy + p 2 (r 2 + 2x 2 )

[0051] y' = y + (y - c y )(k 1 r 2 + k 2 r 4 ) + p 1 (r 2 + 2y 2 ) + 2p 2 xy

[0052] where k 1 , k 2 is the radial distortion coefficient, p 1 , p 2 is the tangential distortion coefficient, and r 2 = x 2 + y 2 is the square of the distance to the optical axis;

[0053] Optimize by the least squares method and the maximum likelihood method. Use the least squares method to solve the preliminary distortion coefficients k 1 , k 2 , p 1 , p 2 , and minimize the following error function:

[0054]

[0055] where x ij is the pixel coordinate in the ideal undistorted case, and x' ij is the pixel coordinate after actual distortion; Further optimize the estimated distortion parameters using the maximum likelihood method to obtain more accurate internal and external parameters, and solve the parameter values according to the pixel coordinates and physical coordinates of the detected feature points.

[0056] As a preferred solution of the machine vision-based target positioning system described in the present invention, wherein:

[0057] Image acquisition module, select appropriate key hardware for machine vision, collect image information meeting the requirements of accuracy and clarity, install multiple visual cameras for aligning tie rods and visual cameras for aligning suspension shafts, and accurately detect the deviation distances of the two tie rods and the suspension shaft;

[0058] Image preprocessing module, preprocess the collected images to improve the image quality and the accuracy of feature extraction; The preprocessing includes denoising, grayscale conversion, edge enhancement, feature point detection, and feature point extraction;

[0059] Coordinate system establishment and internal and external parameter matrix generation module, establish the pixel coordinate system and the image coordinate system, establish the camera coordinate system and the world coordinate system, realize coordinate transformation, and accurately map the target position in the image to the actual space position;

[0060] An automatic calibration module uses a calibration board for camera calibration, calculates the internal and external parameter matrices through image acquisition at multiple angles and distances, thereby correcting the geometric distortion of the camera to ensure accurate camera imaging; the system is controlled by a PLC, processes and feeds back the received visual detection data to achieve automatic calibration and positioning.

[0061] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of a target positioning method based on machine vision.

[0062] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of a target positioning method based on machine vision.

[0063] Advantages of the present invention: The SIFT algorithm adopted in one solution of the target positioning method based on machine vision provided by the present invention has strong robustness to scale changes, rotations, illumination changes, and certain degrees of perspective changes of images, and is applicable to complex environments in industrial scenarios. By establishing the conversion from the pixel coordinate system, the image coordinate system, and the normalized camera coordinate to the camera coordinate system, the internal and external parameter matrices are generated, and combined with the camera calibration method, the problem of converting the image coordinate to the world coordinate is solved. During the calibration process, the checkerboard calibration board and the least squares method are used to improve the calculation accuracy of the internal and external parameters, effectively reduce the calibration error, and improve the positioning accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0065] Figure 1 It is the overall flowchart of a target positioning method based on machine vision provided by the first embodiment of the present invention.

[0066] Figure 2 It is the layout diagram of the camera device of a target positioning method based on machine vision provided by the first embodiment of the present invention.

[0067] Figure 3 It is the schematic diagram of the coordinate conversion between the camera coordinate system and the image coordinate system of a target positioning method based on machine vision provided by the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0069] Example 1, referring to Figures 1 to 3 , which is an embodiment of the present invention, provides a target positioning method based on machine vision, including:

[0070] S1: Select key hardware for machine vision and collect image information.

[0071] Further, the selection of key hardware for machine vision includes selecting a light source and a lighting method, and determining the resolution of the camera and the lens.

[0072] Furthermore, for lens selection, the main consideration factor is the focal length; assuming the size of the field of view to be FOV, with the horizontal and vertical directions being FOV H and FOV V , the horizontal and vertical sizes of the camera chip target surface are S H and S V , the detection distance is d w , then the horizontal component calculation formula and vertical component calculation formula of the focal length f are expressed as:

[0073]

[0074] where f H represents the horizontal focal length size, and f V represents the vertical focal length size.

[0075] Furthermore, as Figure 2 shown, the collection of image information includes using two cameras to collect images; the first camera collects the images of the round hole surface of the pull rod and the end face of the rear suspension shaft of the pull rod from the front to detect the alignment of the holes of the pull rod and the alignment of the pull rod and the suspension shaft; the second camera collects the side images of the pull rod and the suspension shaft to detect the hole alignment of the pull rod and the penetration degree of the suspension shaft pin.

[0076] It should be noted that the specific method of collecting images is to support the pull plate on the locking beam, drive the axis shifting device to move and approach the locking beam; fix the axis shifting device and the locking beam with a locking pin; place the hoisting device at the front shaft storage slot, gradually lift the lifting platform, and align the shaft to be installed with the shaft hole; telescope the transverse cylinder to drive the sliding table and the electric push rod and the hanging shaft on the sliding table to move horizontally, and at the same time fine-tune the up and down position of the platform to make the shaft to be installed and the shaft hole completely aligned; the axis shifting transfer device moves the pin shaft from the storage position to the middle pinning position through the clamp; start the electric push rod, connect the electric push rod and the shaft to be installed, and gradually advance the hanging shaft until the pin is in place; remove the electric push rod and the hanging shaft connection device, retract the electric push rod, and lower the lifting platform to the lowest level; remove the locking pin of the axis shifting device and the locking beam, and the axis shifting device retreats and distances itself from the locking beam.

[0077] S2: Perform image preprocessing and use the SIFT algorithm to automatically detect the key points and feature points of the image.

[0078] Furthermore, the image preprocessing includes grayscale processing based on color features, image filtering, and binarization of the image to achieve rough ROI positioning.

[0079] Furthermore, a suitable grayscale transformation method is selected to transform the three-channel color RGB image into a single-channel grayscale image with the same value range; the grayscale value at the pixel point P(x,y) after the grayscale transformation is set to Gray(x,y).

[0080] Furthermore, the average grayscale processing method takes the average value of the RGB three-channel component values ​​as the grayscale value formula of the corresponding pixel after grayscale transformation as follows:

[0081]

[0082] Furthermore, the weighted average grayscale processing method is used to generalize the average grayscale processing method. The grayscale values ​​of the pixels corresponding to the grayscale transformation are calculated after the RGB three-channel component values ​​are assigned appropriate weights. The formula is expressed as follows:

[0083] Gray(x,y)=r 1 R(x,y)+r 2 G(x,y)+r 3 B(x,y)

[0084] Among them, R(x,y) represents the pixel point of the R channel, G(x,y) represents the pixel point of the G channel, and B(x,y) represents the pixel point of the B channel; 1 、r 2 、r 3 Represents the initial weight assigned to each channel component.

[0085] Further, the rough ROI positioning process is based on the grayscale image after the smoothing process in the previous section, and performs inverse binary threshold processing. During the processing, a threshold thresh is selected, and the formula is expressed as:

[0086]

[0087] Further, obtain the center point of the largest connected component. It can be seen from the binary image that the position of the largest connected component is the position where the target to be detected is located, and its center point is also approximately the center point of the area. After obtaining the center point position, rough ROI positioning can be performed according to the round hole size information.

[0088] Further, obtain the rough ROI positioning area. Based on the approximate position of the target center point obtained above, according to the size information, a rectangular frame obtained by expanding several pixels from the center point to the surrounding is the rough positioning area containing the target to be detected.

[0089] Perform cropping on the rough ROI positioning area, and only process the rough positioning area in the subsequent processing.

[0090] Further, the automatic detection of key points and feature points of the image using the SIFT algorithm includes performing scale-space extreme value detection, searching for image positions at all scales, and identifying potential scale- and rotation-invariant interest points through Gaussian differential functions.

[0091] Statistically analyze all initially detected key points and calculate the total number of key points.

[0092] If the number of key points far exceeds the preset threshold range, adjust the Gaussian blur parameter to reduce redundant key points. If the number of key points is too small, it may be necessary to relax the detection criteria to obtain more feature points.

[0093] Perform spatial analysis on the distribution of key points in the image and calculate the distribution of key points within the image area. Grid division or clustering algorithms can be used to evaluate the uniformity of the distribution. The formula for the distribution uniformity analysis is to divide the image into m×n grid regions, and the number of key points in each grid region is N ij , calculate the uniformity of the key point distribution. Use variance to measure the uniformity of the distribution:

[0094]

[0095] Among them, μ is the average number of key points per grid, and the calculation formula is:

[0096]

[0097] If the variance exceeds the set threshold, it means that the distribution is uneven and feedback adjustment is required.

[0098] Adjust the Gaussian blur scale. If the number of key points is small or unevenly distributed, the detection can be refined by increasing the number of Gaussian blur scale levels, allowing more features to be recognized. If the number of key points is too large and redundant, the Gaussian blur scale can be decreased or the interval between scales can be increased to reduce the detection of details and prevent excessive redundant features.

[0099] Reconstruct the scale space and detect key points again according to the new Gaussian blur parameters and threshold settings.

[0100] Re - count and analyze the newly detected key points to verify whether the adjustment has achieved the expected effect. If the result is still not ideal, continue to trigger the feedback adjustment and further adjust the parameters.

[0101] Iterate the feedback formula. In the i - th feedback iteration, assume that the number of detected key points is The ideal range of the number of key points is [ N min ,N max] . The feedback adjustment formula is expressed as:

[0102]

[0103] The iteration termination condition can be set to reach the target range of the number of key points or exceed the maximum number of iterations iter max :

[0104] Stop if or i≥iter max

[0105] Let the target range of the number of key points be [ N min ,N max] , and the feedback condition is:

[0106] N min ≤N keypoints ≤N max

[0107] Distribution uniformity condition: Let the distribution variance threshold be The feedback condition is:

[0108]

[0109] Perform precise key point localization. At each candidate position, determine the position and scale by fitting a refined model, and the key points are selected based on their stability.

[0110] Determine the main direction of the key points. Based on the gradient direction of the local image, assign one or more directions to each key point position. All subsequent operations on the image data are transformed relative to the direction, scale, and position of the key points, so as to provide invariance to these transformations;

[0111] Generate SIFT feature vectors. Measure the local gradient of the image at a selected scale within the neighborhood around each key point.

[0112] It should be noted that the present invention uses an industrial camera to collect images between the drawbar and the hanging shaft, and then transmits them to an industrial control computer. Feature point detection, description, and edge detection are performed through the OpenCV vision algorithm to obtain the spatial position and edges of the shaft hole. The conversion between different vision coordinate systems and the robotic arm coordinate system is achieved through camera calibration. Finally, by transmitting the visual information to the industrial control center, the multi-degree-of-freedom platform is controlled by a PLC to perform shaft alignment work, completing the automatic and rapid alignment between the drawbar, the hanging shaft, and the drawbar.

[0113] S3: Establish a coordinate system to generate an internal parameter matrix, establish a camera coordinate system to generate an external parameter matrix, and perform coordinate transformation by combining the internal and external parameter matrices.

[0114] Further, the establishment of the coordinate system to generate the internal parameter matrix includes that the pixel coordinate system is the coordinate of the image with respect to the determined target point, a coordinate system with the basic unit of pixel, established on the photosensitive element of the camera. There is a point (u 0 , v 0 ) in the pixel coordinate system.

[0115] The image coordinate system is a coordinate system with physical length as the unit, also established on the photosensitive element of the camera. Then the image coordinates are (u, v):

[0116]

[0117] Convert the image coordinates (u, v) to the normalized camera coordinates (x n , y n ), and the formula is expressed as:

[0118]

[0119] Among them, f x and f y represent the pixel units of the camera focal length; c x and c y represent the pixel coordinates of the principal point of the image;

[0120] Convert the normalized camera coordinates to the camera coordinate system. Map the normalized camera coordinates into the camera coordinate system. Assuming that the depth Z of the object c is known, then the formula is expressed as:

[0121] X c = x n ·Z c , Y c = y n ·Z c , Z c = Z c

[0122] The camera coordinate system is a reference system based on the camera to describe the position of the target, with physical length as the unit. The directions of the X-axis and Y-axis are the same as those of the x-axis and y-axis respectively. The Z-axis is perpendicular to the camera's photosensitive element, and the direction away from the photosensitive element is taken as positive. The center of the lens is the origin of this coordinate system. Then, the coordinates of a point P(x, y) in the camera coordinate system are p c (x c , y c , z c ).

[0123] As Figure 3 shown, if there is a point in the camera coordinate system and its coordinates in the image coordinate system are, according to the similarity theorem, the following formula can be obtained:

[0124]

[0125] It can be found from the above formula that there is a negative sign relationship in the coordinate transformation. Therefore, the imaging plane can be transformed, as Figure 3 shown by the dashed plane. At this time, the transformation relationship is as follows:

[0126]

[0127] For the convenience of calculation, the above formula is transformed into the homogeneous coordinate matrix form as follows:

[0128]

[0129] Its inverse transformation is as follows:

[0130]

[0131] If the relationships in the above formulas are substituted and sorted out, it is as follows:

[0132]

[0133] In the vision system, by introducing the world coordinate system, only the position of the target in the world coordinate system needs to be known, and the external parameter matrix can be obtained through the corresponding formula. Assuming that there is a point in the world coordinate system with corresponding coordinates, the relationship between the two is as follows:

[0134] p c = Rp w + T

[0135] Among them, R represents the rotation matrix, T represents the translation amount, and p w represents the coordinates of the point in the world coordinate system;

[0136]

[0137] T = (t x , t y , t z ) T

[0138] For the convenience of calculation, the above three formulas can be jointly expressed in the form of a homogeneous coordinate matrix as shown in the following formula.

[0139]

[0140] Among them, determining the pose of the camera coordinate system is called the external parameter, that is, the pose of the camera, and the matrix M is called the external parameter matrix, and its inverse transformation is shown in the following formula.

[0141]

[0142] Integrating the whole process, the image coordinates (u, v) are converted into the world coordinate system through the internal and external parameters.

[0143] It should be noted that the pixel coordinate system is a coordinate system for determining the coordinates of the target point, with the basic unit being pixels, and it is established on the photosensitive element of the camera. The horizontal direction is defined as the Opu axis, with the right direction being positive, the vertical direction is defined as the Opv axis, with the downward direction being positive, and the upper left corner of the image is defined as the coordinate system origin Op. The image coordinate system is a coordinate system for determining the coordinates of different points, with the physical length as the unit, and it is also established on the photosensitive element of the camera. The directions of the Oix axis and Oiy axis are the same as those of the Opu axis and Opv axis respectively, and the projection point of the optical axis on the photosensitive element is defined as the coordinate system origin Oi, which represents the position of the image coordinate system origin in the pixel coordinate system where it is located, generally 1 / 2 of the image pixel size.

[0144] S4: Through the method of camera calibration, target positioning calibration is carried out.

[0145] Furthermore, the calibration by combining the internal and external parameter matrices includes performing camera calibration, selecting a suitable checkerboard calibration board for printing, and determining the physical size of the relevant points on the calibration board to ensure the accuracy of the calibration board.

[0146] Fix the calibration board on a certain flat plate, and within the camera's field of view, change the position and angle of the calibration board and take multiple shots correspondingly.

[0147] Use the algorithm to detect the grid corner feature points of the checkerboard calibration board in the image.

[0148] Distortion coefficient solution and optimization. After initially solving the internal and external parameters of the camera, the image distortion factor is considered. Establish radial distortion and tangential distortion models:

[0149] x' = x + (x - c x )(k 1 r 2 + k 2 r 4 ) + 2p 1 xt + p 2 (r 2 + 2x 2 )

[0150] y' = y + (y - c y )(k 1 r 2 + k 2 r 4 ) + p 1 (r 2 + 2y 2 ) + 2p 2 xy

[0151] where k 1 , k 2 are the radial distortion coefficients, and p 1 , p 2 are the tangential distortion coefficients, and r 2 = x 2 + y 2 is the square distance to the optical axis.

[0152] Optimization using the least squares method and the maximum likelihood method. Use the least squares method to solve the initial distortion coefficients k 1 , k 2 , p 1 , p 2 , and minimize the following error function:

[0153]

[0154] where x ij is the pixel coordinate in the ideal undistorted case, and x' ij is the pixel coordinate after actual distortion. Use the maximum likelihood method to further optimize the estimated distortion parameters to obtain more accurate internal and external parameters, and solve the parameter values based on the pixel coordinates and physical coordinates of the detected feature points.

[0155] It should be noted that it is an improved traditional calibration method. It not only has the advantage of low error of the camera self-calibration method but also does not have the disadvantages of cumbersome process and high requirements of other traditional calibration methods. It combines the characteristics of traditional methods and self-calibration methods.

[0156] Embodiment 2, an embodiment of the present invention, provides a target positioning method based on machine vision. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0157] In this embodiment, the parts alignment and detection in an automated assembly production line are used as the test object, and the purpose is to achieve high-precision target position detection and alignment through the target positioning method based on machine vision. The experimental system includes two high-resolution industrial cameras (Camera 1 and Camera 2), a variety of optional lenses, an LED light source, and a control computer. The system hardware selection is based on conditions such as the size of the target part, surface characteristics, and ambient light. Camera 1 is used to capture the overall contour of the part from the front, and Camera 2 is used to capture key feature points from the side. The focal length of the lens selection is determined by calculating the field of view size and the detection distance to ensure the accuracy and clarity of the captured images.

[0158] In the experiment, image acquisition of the selected part is first performed. During the image preprocessing, the captured color image is grayscale processed, filtered to remove noise, and the target detection area is further determined through ROI (Region of Interest) rough positioning. In the preprocessed image, the SIFT (Scale-Invariant Feature Transform) algorithm is used to extract key points and feature points. The extraction process of feature points includes steps such as scale space extreme value detection, precise key point positioning, and main direction assignment to ensure stability under different perspectives and lighting conditions. Subsequently, the system establishes a pixel coordinate system and an image coordinate system, and generates the corresponding internal parameter matrix to describe the imaging characteristics of the camera. The relationship between the camera coordinate system and the world coordinate system is described by the external parameter matrix, and the precise mapping from image coordinates to camera coordinates is achieved by combining the internal parameter matrix.

[0159] Finally, camera calibration is performed by using a standard checkerboard calibration board. After taking calibration images at multiple different angles, the pixel coordinates and known physical coordinates of the checkerboard grid points are detected by the algorithm, and the internal and external parameter matrices are generated and optimized. Through the error correction of the least squares method and the maximum likelihood method, the high-precision and low-error performance of the system in target positioning are ensured. Part of the results is shown in Table 1.

[0160] Table 1 Data Recording Form

[0161]

[0162]

[0163] As can be seen from the tabular data, through this embodiment, the object localization method based on machine vision achieved a high success rate of key point matching and a low localization error. Compared with the existing systems based on simple feature extraction algorithms, this embodiment demonstrated extremely high stability and robustness on different experimental objects. In the data, the localization error of part C was 0.4 mm, and the success rate of key point matching reached 96%. In contrast, traditional systems could only achieve a matching rate of about 90% and a localization error of 0.7 - 1.0 mm in similar scenarios, which demonstrated the advantages of using the SIFT algorithm in complex environments. With the support of high-precision hardware selection, the imaging clarity index of each part generally remained at a high level, ensuring the stability of image acquisition.

[0164] Embodiment 3, an embodiment of the present invention, provides an object localization system based on machine vision, including:

[0165] An image acquisition module, which selects appropriate key machine vision hardware to acquire image information meeting accuracy and clarity requirements, installs multiple visual cameras for aligning pull rods and visual cameras for aligning the hanging shaft holes, and precisely detects the deviation distances between the two pull rods and the hanging shaft.

[0166] An image preprocessing module, which preprocesses the acquired images to improve the image quality and the accuracy of feature extraction; the preprocessing includes denoising, grayscale conversion, edge enhancement, feature point detection, and extraction of feature points.

[0167] A coordinate system establishment and internal and external parameter matrix generation module, which establishes the pixel coordinate system and the image coordinate system, and the camera coordinate system and the world coordinate system, realizes coordinate transformation, and precisely maps the target position in the image to the actual space position.

[0168] An automatic calibration module, which uses a calibration board for camera calibration, calculates the internal and external parameter matrices through image acquisition at multiple angles and distances, thereby correcting the geometric distortion of the camera to ensure accurate camera imaging; the system is controlled by a PLC, processes and feeds back the received visual detection data, and realizes automatic calibration and localization.

[0169] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0170] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0171] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROMs). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0172] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

[0173] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A target positioning method based on machine vision, characterized in that: include: Select key machine vision hardware and collect image information; Perform image preprocessing and use SIFT algorithm to automatically detect key points and feature points of the image; Establish a coordinate system to generate an internal parameter matrix, establish a camera coordinate system to generate an external parameter matrix, and combine the internal and external parameter matrices to perform coordinate transformation; The target positioning is calibrated by camera calibration method.

2. The target positioning method based on machine vision according to claim 1, characterized in that: The selection of key hardware for machine vision includes selecting the light source and lighting method, and determining the resolution and lens of the camera; For lens selection, the main consideration is focal length; let the field of view of the shot be FOV, and the horizontal and vertical directions are FOV H and FOV V The horizontal and vertical sizes of the camera chip target surface are S H and S V , the detection distance is d w , then the calculation formulas for the horizontal and vertical components of the focal length f are expressed as: Among them, f H Indicates the horizontal focal length, f V Indicates the vertical focal length.

3. The target positioning method based on machine vision as claimed in claim 2, characterized in that: The image information collection includes using two cameras to collect images; camera No. 1 collects the end face images of the circular hole surface of the tie rod and the rear suspension shaft of the tie rod from the front, and detects the alignment of the tie rod with the hole and the alignment of the tie rod with the suspension shaft; camera No. 2 collects the side images of the tie rod and the suspension shaft, and detects the alignment of the tie rod with the hole and the degree of pinning of the suspension shaft.

4. The target positioning method based on machine vision as claimed in claim 3, characterized in that: The image preprocessing includes grayscale processing based on color features, image filtering, and binarization of the image to achieve ROI rough positioning; Select a suitable grayscale transformation method to transform the three-channel color RGB image into a single-channel grayscale image with the same value range; set the grayscale value at the pixel point P(x,y) after grayscale transformation to Gray(x,y); The average grayscale processing method takes the average value of the RGB three-channel component values ​​as the grayscale value of the corresponding pixel after grayscale transformation. The formula is expressed as: The weighted average grayscale processing method is used to generalize the average grayscale processing method. The grayscale values ​​of the pixels corresponding to the grayscale transformation are calculated after the RGB three-channel component values ​​are assigned appropriate weights. The formula is expressed as follows: Gray(x,y)=r1R(x,y)+r2G(x,y)+r3B(x,y) Among them, R(x,y) represents the pixel point of the R channel, G(x,y) represents the pixel point of the G channel, and B(x,y) represents the pixel point of the B channel; r1, r2, and r3 represent the initial weights assigned to each channel component; The ROI rough positioning process is based on the grayscale image after smoothing in the previous section, and performs anti-binarization threshold processing. During the processing, the threshold thresh is selected, and the formula is expressed as: Among them, src(x,y) represents the pixel value at coordinates (x,y) in the input image; dst(x,y) represents the pixel value at coordinates (x,y) in the output image; maxval represents the maximum value to be assigned to those pixels that meet the conditions; Get the center point of the maximum connected domain. From the binary image, it can be seen that the position of the maximum connected domain is the location of the target to be detected, and its center point is also the approximate center point of the area. After obtaining the center point position, the ROI can be roughly positioned according to the circular hole size information; Obtain the ROI rough positioning area. Based on the approximate position of the target center point obtained above, a rectangular frame obtained by extending a number of pixels from the center point to the surrounding area according to the size information is the rough positioning area containing the target to be detected; The ROI coarse positioning area is cropped, and the subsequent processing process only needs to process the coarse positioning area.

5. The target positioning method based on machine vision as claimed in claim 4, characterized in that: The automatic detection of key points and feature points of an image using the SIFT algorithm includes performing scale space extremum detection, searching for image positions at all scales, and identifying potential scale- and rotation-invariant points of interest through Gaussian differential functions; When performing scale space extreme value detection, a dynamic feedback adjustment mechanism is added to monitor the number and distribution of key points detected in real time; if too few or too many key points are detected at a specific scale, the scale parameters of the Gaussian blur are automatically adjusted; Accurately locate key points. At each candidate location, a fine-tuned model is used to determine the position and scale. The key points are selected based on their stability. Determine the main direction of the key points, assign one or more directions to each key point position based on the local gradient direction of the image, and all subsequent operations on the image data are transformed relative to the direction, scale and position of the key points, thereby providing invariance to these transformations; Generate SIFT feature vectors that measure the local gradient of the image at a selected scale in a neighborhood around each keypoint.

6. The target positioning method based on machine vision according to claim 5, characterized in that: The coordinate transformation by combining the intrinsic and extrinsic matrix includes: the pixel coordinate system is the coordinate system of the image with respect to the determined target point, the basic unit is the pixel coordinate system, which is established on the photosensitive element of the camera; the image coordinate system is the coordinate system with the physical length as the unit, which is also established on the photosensitive element of the camera; The image coordinates ( u,v ) Convert to normalized camera coordinates ( x n ,y n) , the formula is: Among them, f x and f y The pixel unit that represents the focal length of the camera; c x and c y Represents the pixel coordinates of the principal point of the image; The normalized camera coordinates are transformed into the camera coordinate system, and the normalized camera coordinates are mapped into the camera coordinate system, assuming that the depth Z of the object c Known, the formula is expressed as: X c =x n ·Z c ,Y c =y n ·Z c ,Z c =Z c The camera coordinate system is a reference system based on the camera as a reference object to describe the position of the target. The axis direction and the axis direction are defined as the same as the axis and the axis respectively. The axis is perpendicular to the camera photosensitive element and is taken away from the photosensitive element as positive. The center of the lens is the origin of the coordinate system; then the coordinate of a point P(x, y) in the camera coordinate system is p c (x c ,y c ,z c ); To introduce the world coordinate system into the visual system, we only need to know the position of the target in the world coordinate system and use the corresponding formula to calculate the extrinsic parameter matrix; assuming that there is a point in the world coordinate system, the corresponding coordinates are, and the relationship between the two is shown in the following formula. p c =Rp w +T Among them, R represents the rotation matrix, T represents the translation, and p w Represents the coordinates of a point in the world coordinate system; Comprehensive the whole process, image coordinates ( u,v ) The internal and external parameters are converted into world coordinates.

7. The target positioning method based on machine vision according to claim 6, characterized in that: The method of camera calibration includes performing camera calibration, selecting a suitable chessboard calibration plate for printing, and determining the physical dimensions of relevant points of the calibration plate to ensure the accuracy of the calibration plate; Fix the calibration plate on a plane, change the position and angle of the calibration plate within the camera field of view, and take multiple shots accordingly; The algorithm is used to detect the grid corner feature points of the chessboard calibration plate in the image; Distortion coefficient solution and optimization: After initially solving the camera internal and external parameters, consider the image distortion factor; establish radial distortion and tangential distortion models: x'=x+(x-c x )(k1r 2 +k2r 4 )+2p1xy+p2(r 2 +2x 2 ) y'=y+(y-c y )(k1r 2 +k2r 4 )+p1(r 2 +2y 2 )+2p2xy Among them, k1, k2 are radial distortion coefficients, p1, p2 are tangential distortion coefficients, r 2 =x 2 +y 2 is the square distance to the optical axis; Least squares and maximum likelihood optimization, use the least squares method to solve the preliminary distortion coefficients k1, k2, p1, p2, and minimize the following error function: Among them, x ij is the pixel coordinate under ideal distortion-free conditions, x' ij is the pixel coordinate after actual distortion; the maximum likelihood method is used to further optimize the estimated distortion parameters to obtain more accurate internal and external parameters, and the parameter value is solved according to the pixel coordinates and physical coordinates of the detected feature points.

8. A system using the machine vision-based target positioning method according to any one of claims 1 to 7, characterized in that: It includes image acquisition module, image preprocessing module, coordinate system establishment and internal and external parameter matrix generation module, and automatic calibration module; Image acquisition module, select appropriate key machine vision hardware, collect image information that meets the requirements of accuracy and clarity, install multiple tie rod alignment vision cameras and suspension shaft hole alignment vision cameras, and accurately detect the deviation distance between the two tie rods and the suspension shaft; Image preprocessing module, which preprocesses the collected images to improve the image quality and the accuracy of feature extraction; Preprocessing includes denoising, graying, edge enhancement, feature point detection, and feature point extraction; Coordinate system establishment and internal and external parameter matrix generation module, pixel coordinate system and image coordinate system establishment, camera coordinate system and world coordinate system establishment, realize coordinate conversion, and accurately map the target position in the image to the actual space position; The automatic calibration module uses a calibration plate to calibrate the camera. It calculates the internal and external parameter matrices through image acquisition at multiple angles and distances, thereby correcting the camera's geometric distortion and ensuring accurate camera imaging. The system is controlled by PLC to process and feedback the received visual inspection data to achieve automatic calibration and positioning.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the machine vision-based target positioning method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the machine vision-based target positioning method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Pipeline change detection method based on deep learning and unmanned aerial vehicle images

    CN111967337A

  • Solid engine multi-cylinder-section butt joint guiding measurement algorithm based on binocular vision

    CN112362034A

  • Parameter optimization camera calibration method

    CN113160333A

  • Hand key point space coordinate acquisition method based on binocular vision

    CN114119739A

  • Visual positioning method and system based on deep learning and storage medium

    CN115272457A

Cited By

  • Fusion calibration method and system for semiconductor measurement equipment

    CN121010651A

  • A fusion calibration method and system for semiconductor measurement equipment

    CN121010651B

  • Visual positioning method, system and device based on three-dimensional model critical dimension and medium

    CN121074130A