A method, system and related device for stitching images taken by multiple cameras for appearance inspection
By using calibration plates and weighting error minimization model in multi-camera systems for calibration, and combining feature point screening methods of distance threshold and quality threshold, the problem of low image stitching accuracy when the view of multiple cameras is not overlapped, and efficient image stitching and feature point extraction are achieved.
Patent Information
- Application Number
- CN202510032856.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In a multi-camera system, when the camera's sight area is not overlapped, the prior art image stitching method cannot directly obtain the spatial relationship between the cameras, resulting in low calibration and image stitching accuracy. At the same time, the existing feature point extraction algorithm will generate a large number of redundant feature points in high-resolution images, causing waste of resources, and the extraction and screening of edge points are not accurate enough, affecting the accuracy of subsequent image registration and stitching.
The main camera and the auxiliary camera are calibrated by the calibration plate combined with the weighted error minimization model to obtain the spatial relationship between each camera. Then, the pixel plane transformation matrix of the auxiliary camera to the main camera is determined according to these spatial relationships, and the image of the auxiliary camera is mapped under the pixel plane of the main camera. At the same time, in the process of feature point extraction and screening, distance threshold and quality threshold are used for screening to reduce redundant feature points and improve the quality and accuracy of feature points. Based on high-quality feature point pair sets, the perspective transformation matrix is calculated, and the stitching image is stitched and fused in combination with the viewing projection position.
It improves the image stitching accuracy and overall quality of multi-camera systems in scenes without overlapping sights, reduces resource waste, enhances the accuracy of feature point extraction and screening, and ensures the accuracy of image registration and stitching.
Smart Images

Figure CN119478049B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, system and related device for stitching images taken by multiple cameras for appearance inspection. Background Art
[0002] At present, multi-camera image stitching and industrial positioning technology have been widely used in the field of appearance inspection, especially in the inspection of large-size or ultra-wide display screens. The industrial camera is connected to the main control computer or image acquisition card through a cable, the calibration plate is used as a reference, the light source is used for uniform illumination, and the mechanical moving platform ensures the precise positioning of the calibration plate or display screen. Each camera takes a partial image of the display screen and reconstructs the complete appearance of the entire display screen through the image stitching algorithm. After the system is started, the calibration module obtains the imaging model of each camera and its conversion relationship relative to the main camera. The image stitching module aligns and fuses the multi-camera images according to the calibration results to generate a complete display screen appearance image for appearance inspection.
[0003] However, in a multi-camera system, when the camera fields of view do not overlap, the existing image stitching methods cannot directly obtain the spatial relationship between the cameras, resulting in low calibration and image stitching accuracy. In addition, the existing feature point extraction algorithm will generate a large number of redundant feature points in high-resolution images, resulting in a waste of resources. At the same time, the extraction and screening of edge points are not accurate enough, resulting in low accuracy in subsequent image registration and stitching. Summary of the invention
[0004] The present application provides a method, system and related device for stitching images taken by multiple cameras for appearance inspection, which are used to improve the overall quality and accuracy of image stitching in scenes with non-overlapping fields of view of multiple cameras.
[0005] In a first aspect, the present application provides a method for stitching images taken by multiple cameras for appearance inspection, comprising:
[0006] The main camera and the auxiliary camera are calibrated respectively by using a calibration plate in combination with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from the world coordinate system to the pixel coordinate system;
[0007] Determine a pixel plane transformation matrix from the auxiliary camera to the main camera according to the first projection matrix and the second projection matrix;
[0008] Acquire initial images taken by the main camera and the auxiliary camera, and transform the initial images taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix to obtain a plurality of images to be stitched under the pixel plane of the main camera;
[0009] Projecting the fields of view of all cameras into the world coordinate system and determining whether the projection positions of the fields of view overlap;
[0010] If so, extracting feature points from the images to be stitched, and screening the feature points by using a distance threshold and a quality threshold to obtain a set of feature point pairs;
[0011] The perspective transformation matrix between the images to be spliced is calculated based on the set of feature point pairs, and the images to be spliced are spliced and fused according to the perspective transformation matrix and the field of view projection position.
[0012] Optionally, determining a pixel plane transformation matrix from the auxiliary camera to the main camera according to the first projection matrix and the second projection matrix includes:
[0013] Determine a relative transformation matrix between the first projection matrix and the second projection matrix;
[0014] Decomposing the relative transformation matrix into a rotation matrix and a translation vector, wherein the rotation matrix is used to describe the rotation relationship between the main camera and the auxiliary camera, and the translation vector is used to describe the translation relationship between the main camera and the auxiliary camera;
[0015] An optimization problem is constructed based on the calibration point pairs used in the calibration phase to solve the rotation matrix and the translation vector, and a pixel plane transformation matrix from the auxiliary camera to the main camera is obtained.
[0016] Optionally, feature points are extracted from the images to be stitched, and the feature points are screened by a distance threshold and a quality threshold to obtain a feature point pair set, including:
[0017] Building an image pyramid based on the images to be stitched, extracting feature points on each layer of the image pyramid using a SURF feature extraction algorithm to obtain a feature point set;
[0018] The feature point set is screened according to the distance threshold and the quality threshold to obtain a feature point pair set, wherein the distance threshold is used to screen and remove feature point pairs whose distance is less than the distance threshold, and the quality threshold is used to screen and remove feature points whose response values are less than the quality threshold.
[0019] Optionally, the calculating the perspective transformation matrix between the images to be stitched based on the set of feature point pairs includes:
[0020] Determining an inlier screening threshold according to the resolution and noise level of the image to be stitched, wherein the inlier screening threshold is used to screen reliable points according to the projection error of the feature point pair;
[0021] In combination with the inlier screening threshold and the preset inlier rate threshold, the perspective transformation matrix between the images to be stitched is calculated using the RANSAC algorithm with adaptive iteration termination and the feature point pair set.
[0022] Optionally, the combining the inlier screening threshold and the preset inlier rate threshold, using the adaptive iteration-terminated RANSAC algorithm and the feature point pair set to calculate the perspective transformation matrix between the images to be stitched, includes:
[0023] Calculate the perspective transformation matrix between the images to be stitched using the RANSAC algorithm and the feature point pair set, and obtain the homography matrix calculated in each round of RANSAC iteration;
[0024] Calculating the projection error of each feature point pair in the feature point pair set according to the homography matrix, and screening out the inlier point set whose projection error is less than the inlier point screening threshold;
[0025] Calculating an inlier rate according to the number of feature points in the inlier set and the number of feature points in the feature point pair set;
[0026] When the inlier rate is greater than a preset inlier rate threshold, the iteration is terminated, and the perspective transformation matrix between the images to be spliced is calculated based on the inlier set.
[0027] Optionally, the stitching and fusing the images to be stitched according to the perspective transformation matrix and the field of view projection position includes:
[0028] Performing perspective transformation on the images to be stitched according to the perspective transformation matrix, and determining overlapping areas between the images to be stitched after the perspective transformation;
[0029] Calculating the target weight by using a Gaussian function according to the coordinates of the pixel points in the overlapping area and the preset weight change amplitude;
[0030] The images to be stitched are stitched according to the projection position of the field of view, and the overlapping areas are fused according to the target weight.
[0031] Optionally, the method further includes:
[0032] If not, the images to be stitched are stitched according to the viewing area projection position.
[0033] A second aspect of the present application provides a system for appearance inspection multi-camera image acquisition and stitching, comprising:
[0034] A calibration unit, used to calibrate the main camera and the auxiliary camera respectively through a calibration plate in combination with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from the world coordinate system to the pixel coordinate system;
[0035] a determining unit, configured to determine a pixel plane transformation matrix from the auxiliary camera to the main camera according to the first projection matrix and the second projection matrix;
[0036] an acquisition unit, configured to acquire initial images taken by the main camera and the auxiliary camera, and transform the initial images taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix, so as to obtain a plurality of images to be stitched under the pixel plane of the main camera;
[0037] A judging unit, used for projecting the fields of view of all cameras into the world coordinate system and judging whether the projection positions of the fields of view overlap;
[0038] an extraction unit, configured to extract feature points from the images to be stitched when the judgment result of the judgment unit is yes, and to screen the feature points by using a distance threshold and a quality threshold to obtain a set of feature point pairs;
[0039] The stitching unit is used to calculate the perspective transformation matrix between the images to be stitched based on the set of feature point pairs, and to stitch and fuse the images to be stitched according to the perspective transformation matrix and the view projection position.
[0040] A third aspect of the present application provides a device for stitching images taken by multiple cameras for appearance inspection, the device comprising:
[0041] Processor, memory, input-output unit, and bus;
[0042] The processor is connected to the memory, the input and output unit, and the bus;
[0043] The memory stores a program, and the processor calls the program to execute the first aspect and any optional method for appearance detection multi-camera image stitching in the first aspect.
[0044] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the method for multi-camera image stitching for appearance detection in the first aspect and any optional item in the first aspect is executed.
[0045] It can be seen from the above technical solutions that this application has the following advantages:
[0046] A unified world coordinate system is established through the calibration plate, and the main camera and auxiliary camera are calibrated separately in combination with the weighted error minimization model. The calibration plate provides precise geometric features and clear world coordinate system coordinates, which can serve as a reliable reference for each camera to determine the spatial relationship. The weighted error minimization model fully considers the credibility of different calibration points, optimizes the solution of the projection matrix, and greatly improves the calibration accuracy. Even in scenes where the camera fields of view do not overlap, the spatial relationship between the cameras can be accurately obtained, solving the problem that traditional calibration methods cannot be effectively applied in scenes where the camera fields of view do not overlap.
[0047] Based on this spatial relationship, the pixel plane transformation matrix from the auxiliary camera to the main camera can be further determined, and the initial image taken by the auxiliary camera can be mapped to the pixel plane of the main camera to obtain the image to be stitched, ensuring the accuracy of the mapping relationship between the images taken by each camera. In terms of feature point extraction and screening, feature points are screened according to the distance threshold and quality threshold to eliminate redundant points and evenly distribute feature points. Finally, based on the selected high-quality feature point pair set, the perspective transformation matrix is calculated, and the images to be stitched are stitched and fused in combination with the field of view projection position, which greatly improves the overall quality and accuracy of multi-camera image stitching. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solution in the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A schematic flow chart of an embodiment of a method for stitching images taken by multiple cameras for appearance inspection provided in this application;
[0050] Figure 2 A schematic flow chart of another embodiment of the method for multi-camera image stitching for appearance inspection provided by the present application;
[0051] Figure 3 A schematic diagram of the structure of an embodiment of a system for multi-camera image acquisition and stitching for appearance inspection provided by the present application;
[0052] Figure 4 A schematic diagram of the structure of an embodiment of the device for multi-camera image acquisition and stitching for appearance inspection provided in the present application. DETAILED DESCRIPTION
[0053] The present application provides a method, system and related device for stitching images taken by multiple cameras for appearance inspection, which are used to improve the overall quality and accuracy of image stitching in scenes with non-overlapping fields of view of multiple cameras.
[0054] It should be noted that the method for appearance inspection multi-camera image stitching provided in this application can be applied to a terminal or a server. For example, the terminal can be a smart phone or a computer, a tablet computer, a smart TV, a smart watch, a portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of explanation, this application uses the terminal as the execution subject for example.
[0055] See also Figure 1 , Figure 1 An embodiment of a method for stitching images taken by multiple cameras for appearance inspection provided by the present application includes:
[0056] 101. Calibrate the main camera and the auxiliary camera respectively by using a calibration plate combined with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from a world coordinate system to a pixel coordinate system;
[0057] In a multi-camera system, different cameras have different positions and angles, and the images they capture are spatially independent. In order to achieve the stitching of multi-camera images and accurate positioning measurement, a unified coordinate system must be established to associate the imaging information of each camera. As a known and precise reference object, the calibration plate has a certain position and shape in the world coordinate system. By allowing the camera to image it, the relationship between the camera and this unified world coordinate system can be determined. In industrial positioning scenarios, multiple cameras are often required to perform positioning measurements on objects in the same plane. Especially when the fields of view between cameras do not overlap, it is difficult to directly stitch and calibrate the images. Therefore, in this embodiment, it is necessary to establish a unified coordinate system between cameras with the help of a calibration plate, and according to the position of the feature points of the calibration plate in the image, the intrinsic and extrinsic parameters of the camera are used to establish a mapping from the world coordinate system to the pixel coordinate system of the camera. The pixel coordinate system specifically refers to a coordinate system that describes the pixel position in the image captured by the camera, usually with the upper left corner of the image as the origin, the x-axis as the horizontal direction, and the y-axis as the vertical direction.
[0058] Specifically, the calibration plate is placed within the field of view of the camera, and the main camera and the auxiliary camera are allowed to take images of the calibration plate respectively. The calibration of the main camera and the auxiliary camera is realized by combining the weighted error minimization model. There are a series of feature points with known coordinates (in the world coordinate system) on the calibration plate, such as the intersection of the checkerboard calibration plate. and pixel coordinates , the mapping relationship can be expressed as:
[0059]
[0060] in, yes The projection matrix is defined as:
[0061]
[0062] The above formula can be further decomposed into a nonlinear mapping relationship:
[0063]
[0064] By using multiple sets of and Calibrate data points and build a weighted error minimization model:
[0065]
[0066] By minimizing the error function , solve The best parameter value of . Among them, is the weight, and the basis for weighting is the credibility of the calibration point. For example, when the point is at the edge of the field of view, the error may be large, and the weight can be set low. For the main camera and the auxiliary camera, the above operations are performed respectively, and the first projection matrix of the main camera from the world coordinate system to the pixel coordinate system and the second projection matrix of the auxiliary camera from the world coordinate system to the pixel coordinate system can be obtained.
[0067] In this way, both the main camera and the auxiliary camera determine their own imaging parameters and the relationship with other cameras based on the common world coordinate system, providing a unified benchmark for the establishment of subsequent mapping relationships. And solving the projection matrix using the weighted error minimization model can further improve the calibration accuracy. For those calibration points at the edge of the calibration plate or in areas that may be greatly affected by noise, lower weights can be given, so that the impact of these unreliable points can be reduced when solving the projection matrix. In this way, the transformation relationship between each camera and the world coordinate system in a non-overlapping field of view scene can be accurately quantified.
[0068] 102. Determine a pixel plane transformation matrix from the auxiliary camera to the main camera according to the first projection matrix and the second projection matrix;
[0069] The first projection matrix reflects the conversion rule of the main camera from the world coordinate system to its own pixel coordinate system, and the second projection matrix reflects the conversion rule of the auxiliary camera from the world coordinate system to its own pixel coordinate system. Based on these two known projection matrices, combined with the calibration data points used in the calibration phase for optimization, the pixel plane transformation matrix from the auxiliary camera to the main camera can be determined. Through the pixel plane transformation matrix, in a non-overlapping field of view scene, even if there is no common shooting area between the cameras, the image pixels captured by the auxiliary camera can be accurately mapped to the pixel plane of the main camera according to the pixel plane transformation matrix. Because the pixel plane transformation matrix is derived based on the first projection matrix and the second projection matrix accurately calibrated in step 101, it contains the spatial geometric relationship between the cameras, such as rotation and translation, and the internal parameter information related to imaging, which can ensure the accuracy of the mapping relationship between the images of each camera at the pixel level in the non-overlapping field of view scene.
[0070] The purpose of the pixel plane transformation matrix is to enable the pixel points in the image taken by the auxiliary camera to be accurately converted to the pixel plane of the main camera according to the rules determined by this transformation matrix, so as to realize the processing of multi-camera images in a unified coordinate frame, providing convenience for subsequent image stitching.
[0071] It should be noted that non-overlapping fields of view are not a necessary condition for the technical solution of this application, but one of the specific scenarios to which the technical solution of this application is applicable. If the camera fields of view overlap, the calibration process and subsequent stitching process of this application are still applicable, and can further improve the stitching accuracy and efficiency.
[0072] 103. Acquire initial images taken by the main camera and the auxiliary camera, and transform the initial image taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix to obtain a plurality of images to be stitched on the pixel plane of the main camera;
[0073] Get the initial images taken by the main camera and the auxiliary camera. These initial images record the scene information under the corresponding camera perspective, usually different areas of the display screen, representing different local views of the display screen. The initial images contain the pixel points and their corresponding pixel values in the corresponding camera pixel plane coordinate system. Since the image taken by each camera is obtained based on different camera intrinsic and extrinsic parameters, the position and angle on the pixel plane are different. In order to stitch images from different cameras together, it is necessary to ensure that these images are in the same reference frame, that is, the pixel plane of the main camera. Based on this, the initial image taken by the main camera does not need to be processed, but the initial image taken by the auxiliary camera needs to be processed by coordinate conversion.
[0074] Specifically, for each pixel point in the initial image taken by the auxiliary camera, the coordinate transformation operation is performed according to the matrix multiplication rule through the pixel plane transformation matrix that has been obtained, and the new coordinate obtained is the coordinate position corresponding to the pixel point in the pixel plane of the main camera. Such coordinate transformation operations are performed on the initial images taken by all auxiliary cameras in turn, and the transformation process of multiple images from the auxiliary camera pixel plane to the main camera pixel plane is completed. The images taken by different cameras are aligned to the same reference system, and several images to be stitched are obtained in the pixel plane of the main camera.
[0075] 104. Project the fields of view of all cameras into the world coordinate system, and determine whether the projection positions of the fields of view overlap;
[0076] The field of view (field of view boundary) of each camera is determined by its intrinsic and extrinsic parameters. The intrinsic parameters include the focal length and focus of the camera, and the extrinsic parameters include the position and direction of the camera. The field of view of the camera specifically refers to the spatial angle and area range that the camera can capture. It is determined by factors such as the camera's lens characteristics, sensor size, and installation position and angle. Generally speaking, the field of view is a planar area formed in a three-dimensional space, and this area can be mapped to the world coordinate system through the projection matrix. Based on the projection matrix (first projection matrix and second projection matrix) of a given camera and the image resolution, the projection position of the boundary of each camera's field of view in the world coordinate system can be determined.
[0077] After the fields of view of all cameras are projected into the world coordinate system, it is necessary to check whether there is overlap in the projection positions of these fields of view. The overlap judgment can be completed by calculating the intersection of the geometric areas of the projection positions of the camera fields of view. If the projection positions of the fields of view of two cameras intersect in the world coordinate system, it means that their fields of view overlap. At this time, in order to ensure the stitching effect, step 105 needs to be executed for further stitching processing.
[0078] 105. Extract feature points from the images to be stitched, and screen the feature points by using a distance threshold and a quality threshold to obtain a set of feature point pairs;
[0079] The feature point extraction algorithm is used to detect the same feature points in the images to be stitched. Feature points specifically refer to points with significant information in the images to be stitched. These points are relatively stable under different viewing angles and are easy to match in multiple images. However, in multi-camera stitching, not all feature points can provide effective registration information, especially in high-resolution images. A large number of redundant feature points will be generated during feature point extraction, resulting in a waste of resources. At the same time, the extraction and screening of edge points are not accurate enough, which will also affect the accuracy of subsequent image registration and stitching.
[0080] Therefore, in this embodiment, it is necessary to screen the extracted feature points to ensure that the matched feature points have high quality and accuracy. The screening process depends on the distance threshold and the quality threshold, where the distance threshold is to calculate the spatial distance between all the extracted feature points (such as the commonly used Euclidean distance). If the distance between two feature points is less than the set distance threshold, it means that they may be repeated extractions of the same local feature or redundant points that are too close and affected by noise, and one of them (or both according to specific rules) will be screened out. For example, in an area with a relatively simple image texture, many feature points with close distances may be densely detected. This type of redundancy can be removed by the distance threshold, so that the feature points are more reasonably and evenly distributed in space. The quality threshold is mainly based on the quality attributes of the feature points themselves. For each feature point, its quality-related indicators are calculated. If the quality is lower than the set quality threshold, it means that the feature point may not be stable enough and the recognition is not high in subsequent image matching operations, and it is easy to be confused with other points, so it will be screened out. The quality threshold can be adjusted adaptively according to the image resolution and the density of feature points. For example, in images with high resolution and dense feature points, the quality threshold can be appropriately increased to more strictly screen out high-quality feature points. The remaining feature point pairs after screening form a feature point pair set. These feature points have a high degree of matching in the two images to be stitched and have strong descriptive capabilities, which are suitable for subsequent image registration and stitching.
[0081] 106. Calculate the perspective transformation matrix between the images to be stitched based on the set of feature point pairs, and stitch and fuse the images to be stitched according to the perspective transformation matrix and the view projection position.
[0082] Since the images to be stitched are taken from different angles, there will be differences in geometric deformation, such as changes in image scaling, rotation, and perspective caused by different viewing angles and shooting positions. The perspective transformation matrix can accurately describe the geometric relationship between these images. The perspective transformation matrix can be calculated through a set of feature point pairs, so that the corresponding parts of the pixels in the image can be accurately aligned during stitching to achieve seamless stitching. Although the perspective transformation matrix can handle the problem of geometric alignment between images, relying solely on it may ignore the rationality of the overall stitching layout and boundary parts of the image. Therefore, the stitching fusion can be combined with the field of view projection position, which can control the scope and layout of the stitching from a macro perspective. For example, when processing image stitching with overlapping areas, the perspective transformation matrix mainly focuses on the geometric alignment of images in the overlapping area, while the field of view projection position information can help determine whether the stitching boundary of the image is reasonable, avoiding the situation where the non-overlapping part of a certain image exceeds the range it should actually cover, or does not conform to the spatial layout actually shot by the camera after overall stitching, so that the final stitched complete image is more accurate as a whole and meets the requirements of the actual scene.
[0083] In this embodiment, a unified world coordinate system is established through the calibration plate, and the main camera and the auxiliary camera are calibrated separately in combination with the weighted error minimization model. The calibration plate provides precise geometric features and clear world coordinate system coordinates, which can be a reliable reference for each camera to determine the spatial relationship. The weighted error minimization model fully considers the credibility of different calibration points, optimizes the solution of the projection matrix, and greatly improves the calibration accuracy. Even in the scene where the camera field of view does not overlap, the spatial relationship between the cameras can be accurately obtained, which solves the problem that the traditional calibration method cannot be effectively applied in the scene where the camera field of view does not overlap. Based on this spatial relationship, the pixel plane transformation matrix from the auxiliary camera to the main camera can be further determined, and the initial image taken by the auxiliary camera is mapped to the pixel plane of the main camera to obtain the image to be spliced, ensuring the accuracy of the mapping relationship between the images taken by each camera. In terms of feature point extraction and screening, the feature points are screened according to the distance threshold and the quality threshold to achieve redundant point elimination and feature point distribution uniformity. Finally, based on the selected high-quality feature point pair set, the perspective transformation matrix is calculated, and the images to be stitched are stitched and fused in combination with the field of view projection position, which greatly improves the overall quality and accuracy of multi-camera image stitching.
[0084] The following is a detailed description of the method for stitching images taken by multiple cameras for appearance inspection provided by this application. Figure 2 , Figure 2 Another embodiment of the method for stitching images taken by multiple cameras for appearance inspection provided by the present application includes:
[0085] 201. Calibrate the main camera and the auxiliary camera respectively by using a calibration plate combined with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from a world coordinate system to a pixel coordinate system;
[0086] In this embodiment, step 201 is similar to step 101 in the aforementioned embodiment and will not be described again here.
[0087] 202. Determine a relative transformation matrix between the first projection matrix and the second projection matrix;
[0088] In this embodiment, steps 202 to 204 are the process of determining the pixel plane transformation matrix through the first projection matrix and the second projection matrix. In order to accurately associate and fuse the image taken by the auxiliary camera with the image taken by the main camera, it is necessary to analyze the spatial transformation relationship between the main camera and the auxiliary camera. The first projection matrix describes the transformation of the main camera from the world coordinate system to its pixel coordinate system, and the second projection matrix corresponds to the same transformation process of the auxiliary camera. Therefore, by determining the relative transformation matrix between the first projection matrix and the second projection matrix, the spatial transformation relationship of the auxiliary camera relative to the main camera can be obtained. For a multi-camera system, the first projection matrix obtained by calibrating the main camera and the second projection matrix of the auxiliary camera , the relative transformation relationship between the two can be expressed as:
[0089]
[0090] in, is the relative transformation matrix from the auxiliary camera to the main camera, and The first and second projection matrices from world coordinates to pixel coordinates for the primary and secondary cameras respectively.
[0091] 203. Decompose the relative transformation matrix into a rotation matrix and a translation vector, wherein the rotation matrix is used to describe the rotation relationship between the main camera and the auxiliary camera, and the translation vector is used to describe the translation relationship between the main camera and the auxiliary camera;
[0092] The relative position relationship of the cameras in space can be described from two key perspectives: rotation and translation. The rotation matrix can reflect the angular difference between the main camera and the auxiliary camera, that is, how much the auxiliary camera rotates around which axes relative to the main camera; the translation vector can clearly indicate the position offset of the auxiliary camera in space relative to the main camera, that is, how much distance it moves in the x, y, z and other directions. Decomposing the relative transformation matrix into the rotation matrix and the translation vector can more intuitively and clearly understand the spatial posture and position relationship between cameras. Based on this, the relative transformation matrix It can be further decomposed into a rotation matrix and translation vectors :
[0093]
[0094] in, It is a 3*3 rotation matrix, which is used to describe the rotation relationship between the two camera coordinate systems. It is a 3*1 translation vector, which is used to describe the translation relationship between the two camera coordinate systems.
[0095] 204. Construct an optimization problem based on the calibration point pairs used in the calibration phase to solve the rotation matrix and the translation vector, and obtain a pixel plane transformation matrix from the auxiliary camera to the main camera;
[0096] Although the initial calculation results of the rotation matrix and translation vector are obtained through the previous steps, various errors are inevitable in the camera calibration process, such as measurement error, calibration point extraction error, camera lens distortion and other factors, which will affect the accuracy of the rotation matrix and translation vector, and thus lead to the inaccuracy of the pixel plane transformation matrix constructed based on them. By using the calibration point pairs in the calibration stage to construct the optimization problem, these actual errors can be comprehensively considered, and the rotation matrix and translation vector can be optimized and adjusted to obtain a more accurate pixel plane transformation matrix, thereby improving the accuracy and reliability of subsequent operations such as image stitching and fusion.
[0097] Specifically, multiple groups are used in the calibration phase. Point to point, build optimization problem solution and :
[0098]
[0099] After constructing the error function, you can choose a suitable optimization algorithm to solve the rotation matrix that minimizes the above error function. and translation vector The specific optimization algorithm includes nonlinear least squares method, etc., which is not limited here. and According to the corresponding mathematical relationship (such as the transformation form under homogeneous coordinates), an accurate pixel plane transformation matrix from the auxiliary camera to the main camera can be constructed:
[0100] Combined with the projection relationship of multi-camera calibration, the pixel points of two non-overlapping view cameras and The mapping can be expressed as:
[0101]
[0102] In this embodiment, the relative transformation relationship between the two cameras is dynamically optimized to enhance the reliability of mapping under non-overlapping viewing areas.
[0103] 205. Acquire initial images taken by the main camera and the auxiliary camera, and transform the initial image taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix to obtain a plurality of images to be stitched on the pixel plane of the main camera;
[0104] In this embodiment, step 205 is similar to step 103 in the above embodiment and will not be described again here.
[0105] 206. Project the fields of view of all cameras into the world coordinate system, and determine whether the projection positions of the fields of view overlap;
[0106] It should be noted that the overlap of the fields of view refers to the partial overlap of the physical scene areas actually photographed by the cameras. The overlap of the field of view projection positions refers to whether these projection areas overlap after the camera fields of view are projected into the world coordinate system established by the calibration plate. If there is overlap in the fields of view, their projection ranges will usually also overlap. However, even if there is no overlap in the fields of view, the projection ranges may also overlap due to the projection transformation. Therefore, in this embodiment, it is necessary to determine whether the field of view projection positions overlap.
[0107] In this embodiment, after the fields of view of all cameras are projected into the world coordinate system, it is necessary to check whether there is overlap in the projection positions of these fields of view. If there is overlap, it means that there are also overlapping areas in the corresponding images. At this time, steps 208 to 211 are executed to take corresponding feature matching, fusion and other further processing measures to ensure the stitching quality. If there is no overlap, step 207 is directly executed to stitch in a relatively simple and direct way.
[0108] 207. Stitching the images to be stitched according to the projection position of the field of view;
[0109] When the field of view projection positions of all cameras do not overlap, it means that there is no overlapping area in the images to be stitched. At this time, the images to be stitched located under the pixel plane of the main camera can be stitched directly according to the field of view projection position. The field of view projection position reflects the spatial layout of the shooting range of each camera in the entire scene. Based on these projection position information, it is clear which part of the scene each camera shoots and their relative position relationship in the overall scene. For example, through the field of view projection, it can be determined that camera A shoots the left part of the scene and camera B shoots the right part. Under the pixel plane of the main camera, the image to be stitched corresponding to camera A can be directly placed at the corresponding position on the left in this order and range, and the image of camera B can be placed at the corresponding position on the right. And so on, the images of each camera can be accurately combined to ensure that the stitched image can fully and orderly present the entire scene being shot, which conforms to the actual spatial layout.
[0110] 208. Construct an image pyramid based on the images to be stitched, extract feature points on each layer of the image pyramid using a SURF feature extraction algorithm to obtain a feature point set;
[0111] In a large image with high resolution, if a single-scale feature extraction method is used directly, a large number of feature points will often be detected, many of which may be redundant. For example, in an area with relatively uniform image texture, many similar feature points may appear densely, which are not very helpful for accurately matching images, but will increase the amount of calculation for subsequent processing. For the images to be spliced, it is necessary to extract enough and accurate feature points to ensure the accuracy of subsequent image matching and splicing, and to avoid extracting too many redundant feature points in a large image with high resolution, which will lead to waste of computing resources and low matching efficiency. Therefore, in this embodiment, by performing a series of downsampling operations on the original image to be spliced, image layers of different resolutions are generated to construct an image pyramid. Each layer of the image pyramid represents a scene presentation at a scale. The bottom layer of the image pyramid is usually the original high-resolution image, which contains rich detail information. As the number of layers increases, the image resolution gradually decreases, showing a macroscopic scene structure.
[0112] The SURF feature extraction algorithm is used to extract feature points on each layer of the constructed image pyramid. For each layer of the image, the integral image is used to quickly scan the local area where each pixel in the image is located (according to the set neighborhood range). By comparing the grayscale values of the pixels in the neighborhood, the local extreme points of the grayscale are found, and these extreme points are marked as feature points of this layer. For each detected feature point, the corresponding local descriptor is calculated. A descriptor vector that can uniquely identify the feature point and has a certain anti-interference ability (such as robustness to changes in illumination, rotation, etc.) is generated for subsequent feature point matching between different images. By repeating this operation on each layer of the image, a set of feature points at different scales can be obtained. Through the image pyramid structure, it is ensured that the corresponding feature points can be extracted on images of different scales, ensuring that the feature points can still be detected even if the image is scaled. The extracted feature points are often key details in the image, such as corners, edges, textures, and other easily identifiable local areas.
[0113] Since the process extracts feature points at different scales, although the bottom layer of high resolution is rich in details, some feature points that can be merged or ignored at small scales will not be repeatedly extracted at large scales because they are extracted at multiple scales at the same time, effectively reducing the number of redundant feature points. As the number of layers increases, the image scale becomes smaller, and the extracted feature points focus more on macroscopic and representative structural features, further avoiding the inclusion of too many insignificant detail feature points, thereby controlling the size of the feature point set as a whole and reducing redundancy.
[0114] In addition, the image pyramid provides multi-scale image representation, and the feature points extracted by the SURF algorithm at different scales can reflect the corresponding relationship of the same scene at different scales. Even if the images taken by different cameras have scale changes (such as different image scaling ratios due to different shooting distances and lens focal lengths), the corresponding feature points can be extracted on the corresponding image layers of different scales, so that when performing image matching, the corresponding feature points can be accurately found across scale differences, ensuring that the matching accuracy is not affected by image scaling.
[0115] 209. Screening the feature point set according to the distance threshold and the quality threshold to obtain a feature point pair set, wherein the distance threshold is used to screen and remove feature point pairs whose distance is less than the distance threshold, and the quality threshold is used to screen and remove feature points whose response values are less than the quality threshold;
[0116] In this embodiment, a dynamic feature point screening strategy is adopted to set a distance threshold for the searched redundant points to the feature points. and quality threshold .
[0117] In the process of extracting image feature points, due to factors such as the image's own texture characteristics, noise interference, and the working mechanism of the feature extraction algorithm, some feature points are often too concentrated in spatial positions. For example, in areas where the image texture is relatively simple and highly repetitive, multiple feature points may be densely detected, and the distances between them are very close. In fact, most of these feature point pairs that are too close are repeated descriptions of the same local feature, which are redundant information for subsequent operations such as image matching and stitching. The purpose of setting a distance threshold is to effectively filter out such redundant feature point pairs, so that the feature points that are finally retained can be relatively evenly distributed in space. This not only reduces unnecessary calculations, but also makes subsequent operations such as feature point matching more efficient and accurate, and avoids incorrect matching due to interference from too many similar feature points. At this time, by setting the distance threshold You can filter out feature points that are too concentrated in space, remove point pairs that are too close, and maintain a uniform distribution of feature points:
[0118] Specifically, for the feature point pair ,like , then remove .
[0119] The quality of different feature points in an image is different, and the quality directly affects its reliability and stability in subsequent operations such as image matching. Feature points with poor quality, such as those with large changes in their corresponding local descriptors under different viewing angles and lighting changes, are difficult to accurately match between different images, which can easily lead to matching errors and affect the quality of image stitching. The purpose of setting a quality threshold is to screen out those feature points with high quality that can remain relatively stable and have good recognition under various image transformation conditions, and only retain these high-quality feature points for subsequent operations, thereby improving the accuracy and reliability of the entire image matching and stitching process. That is, through the quality threshold Filter only retained response values Feature points above the threshold:
[0120]
[0121] , then remove ,in are the eigenvalues of the Hessian matrix.
[0122] The quality threshold It can be adaptively adjusted according to the image resolution and the density of feature points. The response value is an indicator used to measure the significance of feature points and is calculated during the SURF feature point detection process.
[0123] 210. Determine an inlier screening threshold value according to the resolution and noise level of the image to be stitched, where the inlier screening threshold value is used to screen reliable points according to the projection error of the feature point pair;
[0124] In this embodiment, step 210 to step 211 is a process of calculating the perspective transformation matrix between images from the matching point pairs by using the improved RANSAC algorithm, which is described in detail below:
[0125] The resolution and noise level of the images to be stitched are key factors affecting the quality of feature point pairs and the accuracy of subsequent perspective transformation matrix calculations. Resolution determines the level of detail that an image can present. High-resolution images are rich in detail, and feature point positioning is theoretically more accurate, so higher matching accuracy is required between feature point pairs. Low-resolution images have fewer details and are relatively blurred, so there is a certain degree of uncertainty in the location of feature points. The noise level reflects the extent to which the image is disturbed by external factors. High noise can cause deviations in information such as feature point coordinates, resulting in large fluctuations in the projection error of feature point pairs.
[0126] In each iteration of the RANSAC algorithm, the point pair projection error is compared with the inlier screening threshold to determine whether the feature point pair is an inlier. Therefore, in this embodiment, the inlier screening threshold needs to be determined based on the resolution and noise level of the image to be stitched. , the inlier screening threshold It is used to distinguish reliable feature point pairs (inner points) from unreliable feature point pairs (outer points) that are affected by resolution and noise and do not conform to the actual perspective transformation relationship.
[0127] 211. Combining the inlier screening threshold and the preset inlier rate threshold, the RANSAC algorithm with adaptive iteration termination and the feature point pair set are used to calculate the perspective transformation matrix between the images to be stitched;
[0128] If the number of iterations of the RANSAC algorithm itself is set improperly, it may lead to low computational efficiency (too many iterations) or inaccurate results (too few iterations, not enough suitable inliers are found). Therefore, in this embodiment, the inlier screening threshold and the preset inlier rate threshold are combined, and an adaptive iteration termination mechanism is introduced to dynamically control the iteration process of the RANSAC algorithm while accurately screening inliers. Specifically, the following steps are included:
[0129] 1. Use the RANSAC algorithm and the feature point pair set to calculate the perspective transformation matrix between the images to be stitched, and obtain the homography matrix calculated in each round of RANSAC iteration;
[0130] At the beginning of each round of RANSAC iteration, a certain number of feature point pairs are first randomly extracted from the feature point pair set as a sample subset. Based on the extracted feature point pair subset, a linear equation system is constructed according to the geometric principle and mathematical model of perspective transformation. Through multiple pairs of such feature point pair coordinate relationships, a sufficient number of linear equations can be listed to form a linear equation system. Then, the linear algebra method (such as matrix inversion, Gaussian elimination, etc.) is used to solve the equation system and obtain the homography matrix of this iteration. The homography matrix is a temporary matrix that describes the perspective transformation relationship between the two images to be stitched.
[0131] 2. Calculate the projection error of each feature point pair in the feature point pair set according to the homography matrix, and filter out the inlier point set whose projection error is less than the inlier point screening threshold;
[0132] After obtaining the homography matrix of this iteration, it is necessary to perform inlier screening on all the remaining feature point pairs in the feature point pair set. Specifically, the projection error of each feature point pair is calculated using the homography matrix just calculated, and the projection error of each feature point pair is compared with the inlier screening threshold. If the projection error is less than the inlier screening threshold, the feature point pair is marked as an inlier.
[0133] For matching point pairs ,like , then keep ;in, is the homography matrix of the current iteration, The value of is determined by the actual image resolution and noise level.
[0134] 3. Calculate the inlier rate based on the number of feature points in the inlier set and the number of feature points in the feature point pair set;
[0135] After a series of screening operations, the inlier rate is calculated based on the number of feature points in the inlier set and the number of feature points in the feature point pair set. .
[0136] 4. When the inlier rate is greater than the preset inlier rate threshold, the iteration is terminated, and the perspective transformation matrix between the images to be stitched is calculated based on the inlier set.
[0137] In this embodiment, the maximum number of iterations of RANSAC is dynamically adjusted. , when the interior point rate Reaching the preset inlier rate threshold When Terminate the iteration early.
[0138] in, Can be determined according to needs, in some optional application scenarios .
[0139] After the iteration is terminated, all feature point pairs marked as inliers are used to calculate the final perspective transformation matrix between the images to be stitched by a suitable optimization method, such as the least squares method. The perspective transformation matrix is used to perform corresponding geometric transformations on the images to be stitched to achieve accurate image alignment and stitching.
[0140] 212. Perform perspective transformation on the images to be stitched according to the perspective transformation matrix, and determine the overlapping areas between the images to be stitched after the perspective transformation;
[0141] In the multi-camera image stitching scenario, due to the different shooting angles, positions and other factors of different cameras, there are differences in perspective deformation between the images to be stitched. Therefore, it is necessary to perform perspective transformation on the images to be stitched according to the perspective transformation matrix. After completing the perspective transformation, the overlapping area is determined by comparing the pixel coordinate ranges of different images to be stitched. For the overlapping areas, it is necessary to apply appropriate fusion strategies to these areas in a targeted manner to ensure that the transition of the stitched images in the overlapping parts is natural and avoid stitching marks.
[0142] 213. Calculate the target weight by using a Gaussian function according to the coordinates of the pixel points in the overlapping area and the preset weight change range;
[0143] In order to achieve a natural transition of the images to be stitched in the overlapping area and avoid stitching traces, a weighted fusion model is introduced for the overlapping area in this embodiment:
[0144] Assume two images and There is an overlapping area , the pixel value of the fused image It can be expressed as:
[0145]
[0146] in is a weight function that satisfies .
[0147] In order to achieve a natural transition, a Gaussian function can be selected to define the target weight:
[0148]
[0149] in is the center point of the overlapping area, To preset the weight change range, determine the weight change range.
[0150] 214. The images to be stitched are stitched according to the projection position of the field of view, and the overlapping areas are fused according to the target weight.
[0151] The perspective transformation matrix is used to make the images to be stitched geometrically aligned after perspective transformation. At this time, the field of view projection position is combined to ensure that each image is placed in the correct position, which is consistent with the area range of each camera during actual shooting, and ensures that the entire stitched image can accurately and orderly present the complete scene being photographed. For overlapping areas, different target weights are assigned to the pixels in the overlapping area, and the pixel values of different images at the corresponding pixels are fused and calculated according to the target weights, so that when the overlapping area transitions from one image to another, the pixel values can change smoothly, improving the visual quality of the stitching effect.
[0152] After completing the multi-camera image stitching, the images of each camera are seamlessly integrated into a unified pixel plane to form a complete target view. This stitching method not only solves the problem of limited field of view of a single camera, but also retains high-resolution details, providing a reliable image basis for subsequent appearance inspection. Through the unified image generated by stitching, the detection algorithm can perform a global analysis of the entire target surface, avoiding missed detection or false detection problems caused by field of view discontinuities or boundary errors.
[0153] Use filtering algorithms (such as Gaussian filtering or median filtering) to remove noise and enhance image quality for the stitched images. Improve the visibility of subtle scratches through edge enhancement or brightness equalization algorithms to ensure that weak features can be effectively detected. Feature extraction uses gradient analysis operators to extract possible scratch edge features. Use texture analysis (such as gray-level co-occurrence matrix) or morphological operations to highlight the appearance defect features. Defect detection and segmentation use image segmentation algorithms to segment the scratch area, and use calibration information to accurately map the scratch position to the world coordinate system for subsequent repair or quality reporting.
[0154] In this embodiment, based on the calibration plate and the projection matrix, the calibration accuracy is optimized by constructing a weighted error minimization model. Especially in the non-overlapping field of view scene, the spatial geometric relationship between the cameras is accurately described through the relative transformation relationship between the main camera and the auxiliary camera, which solves the problem that the traditional calibration method cannot be effectively applied in the non-overlapping scene. And the SURF feature extraction algorithm combined with the pyramid structure is introduced to extract feature points at different scales. The dynamic screening mechanism of the distance threshold and the quality threshold is adopted to achieve the removal of redundant points and the uniform distribution of feature points, while ensuring the high response value of the extracted points, thereby improving the accuracy and efficiency of image matching. In addition, by introducing an adaptive termination mechanism and an inlier screening threshold, the number of iterations is dynamically adjusted, the screening conditions of the inlier are optimized, and the projection error threshold is flexibly determined according to the image resolution and the noise level, which significantly improves the speed and stability of the registration and meets the requirements of high-resolution industrial scenes. For the overlapping area between cameras, a pixel weighted fusion strategy based on the Gaussian weight function is proposed, which can achieve the natural transition of the spliced image, reduce edge traces, and improve the visual continuity and measurement accuracy of the spliced image.
[0155] The following is a detailed description of the multi-camera image acquisition and stitching system for appearance inspection provided by this application. Figure 3 , Figure 3 Another embodiment of the system for multi-camera image acquisition and stitching of appearance inspection provided by the present application includes:
[0156] The calibration unit 301 is used to calibrate the main camera and the auxiliary camera respectively through a calibration plate combined with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from a world coordinate system to a pixel coordinate system;
[0157] A determining unit 302 is used to determine a pixel plane transformation matrix from the auxiliary camera to the main camera according to the first projection matrix and the second projection matrix;
[0158] The acquisition unit 303 is used to acquire the initial images taken by the main camera and the auxiliary camera, and transform the initial images taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix to obtain a plurality of images to be stitched under the pixel plane of the main camera;
[0159] A determination unit 304 is used to project the fields of view of all cameras into the world coordinate system and determine whether the projection positions of the fields of view overlap;
[0160] The extraction unit 305 is used to extract feature points from the images to be stitched when the judgment result of the judgment unit 304 is yes, and screen the feature points by using a distance threshold and a quality threshold to obtain a set of feature point pairs;
[0161] The stitching unit 306 is used to calculate the perspective transformation matrix between the images to be stitched based on the set of feature point pairs, and stitch and fuse the images to be stitched according to the perspective transformation matrix and the view projection position.
[0162] Optionally, the determining unit 302 is specifically configured to:
[0163] Determine a relative transformation matrix between the first projection matrix and the second projection matrix;
[0164] Decompose the relative transformation matrix into a rotation matrix and a translation vector. The rotation matrix is used to describe the rotation relationship between the main camera and the auxiliary camera, and the translation vector is used to describe the translation relationship between the main camera and the auxiliary camera.
[0165] According to the calibration point pairs used in the calibration phase, an optimization problem is constructed to solve the rotation matrix and translation vector, and the pixel plane transformation matrix from the auxiliary camera to the main camera is obtained.
[0166] Optionally, the extraction unit 305 is specifically used for:
[0167] An image pyramid is constructed based on the images to be stitched, and feature points are extracted on each layer of the image pyramid using the SURF feature extraction algorithm to obtain a feature point set;
[0168] The feature point set is screened according to the distance threshold and the quality threshold to obtain a feature point pair set. The distance threshold is used to screen and remove feature point pairs whose distance is less than the distance threshold. The quality threshold is used to screen and remove feature points whose response value is less than the quality threshold.
[0169] Optionally, the splicing unit 306 is specifically used for:
[0170] Determine the inlier screening threshold according to the resolution and noise level of the image to be stitched, and the inlier screening threshold is used to screen reliable points according to the projection error of the feature point pair;
[0171] Combining the inlier screening threshold and the preset inlier rate threshold, the RANSAC algorithm with adaptive iteration termination and the feature point pair set are used to calculate the perspective transformation matrix between the images to be stitched.
[0172] Optionally, the splicing unit 306 is further configured to:
[0173] Use the RANSAC algorithm and the feature point pair set to calculate the perspective transformation matrix between the images to be stitched, and obtain the homography matrix calculated in each round of RANSAC iteration;
[0174] Calculate the projection error of each feature point pair in the feature point pair set according to the homography matrix, and filter out the inlier point set whose projection error is less than the inlier point screening threshold;
[0175] Calculate the inlier rate based on the number of feature points in the inlier set and the number of feature points in the feature point pair set;
[0176] When the inlier rate is greater than a preset inlier rate threshold, the iteration is terminated, and the perspective transformation matrix between the images to be stitched is calculated based on the inlier set.
[0177] Optionally, the splicing unit 306 is further configured to:
[0178] Performing perspective transformation on the images to be stitched according to the perspective transformation matrix, and determining the overlapping areas between the images to be stitched after the perspective transformation;
[0179] According to the coordinates of the pixels in the overlapping area and the preset weight change range, the target weight is calculated by the Gaussian function;
[0180] The images to be stitched are stitched according to the projection position of the viewshed, and the overlapping areas are fused according to the target weight.
[0181] Optionally, the splicing unit 306 is further configured to:
[0182] When the determination result of the determination unit 304 is no, the images to be stitched are stitched according to the viewing area projection position.
[0183] In this embodiment, the functions of each unit are the same as those described above. Figure 1 or Figure 2 The steps in the above code are corresponding to those in the previous section and will not be repeated here.
[0184] This application also provides a device for appearance inspection multi-camera image acquisition and splicing, please refer to Figure 4 , Figure 4 An embodiment of a device for multi-camera image stitching for appearance inspection provided in the present application includes:
[0185] Processor 401, memory 402, input and output unit 403, bus 404;
[0186] The processor 401 is connected to the memory 402, the input and output unit 403 and the bus 404;
[0187] The memory 402 stores a program, and the processor 401 calls the program to execute any of the above-mentioned methods for appearance inspection and multi-camera image stitching.
[0188] The present application also relates to a computer-readable storage medium, on which a program is stored. When the program is run on a computer, the computer executes any of the above-mentioned methods for multi-camera image acquisition and stitching for appearance inspection.
[0189] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0190] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0191] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk, etc. Various media that can store program codes.
Claims
1. A method for stitching images taken by multiple cameras for appearance inspection, characterized in that: The method comprises: The main camera and the auxiliary camera are calibrated respectively by using a calibration plate in combination with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from the world coordinate system to the pixel coordinate system; Determine a relative transformation matrix between the first projection matrix and the second projection matrix; Decomposing the relative transformation matrix into a rotation matrix and a translation vector, wherein the rotation matrix is used to describe the rotation relationship between the main camera and the auxiliary camera, and the translation vector is used to describe the translation relationship between the main camera and the auxiliary camera; According to the calibration point pairs used in the calibration phase, an optimization problem is constructed to solve the rotation matrix and the translation vector, so as to obtain a pixel plane transformation matrix from the auxiliary camera to the main camera; Acquire initial images taken by the main camera and the auxiliary camera, and transform the initial images taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix to obtain a plurality of images to be stitched under the pixel plane of the main camera; Projecting the fields of view of all cameras into the world coordinate system and determining whether the projection positions of the fields of view overlap; If yes, an image pyramid is constructed based on the image to be stitched, and feature points are extracted on each layer of the image pyramid by using a SURF feature extraction algorithm to obtain a feature point set; The feature point set is screened according to a distance threshold and a quality threshold to obtain a feature point pair set, wherein the distance threshold is used to screen and remove feature point pairs whose distance is less than the distance threshold, and the quality threshold is used to screen and remove feature points whose response value is less than the quality threshold; The perspective transformation matrix between the images to be spliced is calculated based on the set of feature point pairs, and the images to be spliced are spliced and fused according to the perspective transformation matrix and the field of view projection position.
2. The method according to claim 1, characterized in that The calculating the perspective transformation matrix between the images to be stitched based on the set of feature point pairs includes: Determining an inlier screening threshold according to the resolution and noise level of the image to be stitched, wherein the inlier screening threshold is used to screen reliable points according to the projection error of the feature point pair; In combination with the inlier screening threshold and the preset inlier rate threshold, the perspective transformation matrix between the images to be stitched is calculated using the RANSAC algorithm with adaptive iteration termination and the feature point pair set.
3. The method according to claim 2, characterized in that The step of combining the inlier screening threshold and the preset inlier rate threshold, using the adaptive iteration-terminated RANSAC algorithm and the feature point pair set to calculate the perspective transformation matrix between the images to be stitched includes: Calculate the perspective transformation matrix between the images to be stitched using the RANSAC algorithm and the feature point pair set, and obtain the homography matrix calculated in each round of RANSAC iteration; Calculating the projection error of each feature point pair in the feature point pair set according to the homography matrix, and screening out the inlier point set whose projection error is less than the inlier point screening threshold; Calculating an inlier rate according to the number of feature points in the inlier set and the number of feature points in the feature point pair set; When the inlier rate is greater than a preset inlier rate threshold, the iteration is terminated, and the perspective transformation matrix between the images to be spliced is calculated based on the inlier set.
4. The method according to claim 1, characterized in that: The step of stitching and fusing the images to be stitched according to the perspective transformation matrix and the field of view projection position comprises: Performing perspective transformation on the images to be stitched according to the perspective transformation matrix, and determining overlapping areas between the images to be stitched after the perspective transformation; Calculating the target weight by using a Gaussian function according to the coordinates of the pixel points in the overlapping area and the preset weight change amplitude; The images to be stitched are stitched according to the projection position of the viewing area, and the overlapping areas are fused according to the target weight.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: If not, the images to be stitched are stitched according to the viewing area projection position.
6. A system for multi-camera image acquisition and splicing for appearance inspection, characterized in that: The system comprises: A calibration unit, used to calibrate the main camera and the auxiliary camera respectively through a calibration plate in combination with a weighted error minimization model to obtain a first projection matrix of the main camera from a world coordinate system to a pixel coordinate system, and a second projection matrix of the auxiliary camera from the world coordinate system to the pixel coordinate system; a determination unit, configured to determine a relative transformation matrix between the first projection matrix and the second projection matrix; decompose the relative transformation matrix into a rotation matrix and a translation vector, wherein the rotation matrix is used to describe a rotation relationship between the main camera and the auxiliary camera, and the translation vector is used to describe a translation relationship between the main camera and the auxiliary camera; and construct an optimization problem based on the calibration point pairs used in the calibration phase to solve the rotation matrix and the translation vector, so as to obtain a pixel plane transformation matrix from the auxiliary camera to the main camera; an acquisition unit, configured to acquire initial images taken by the main camera and the auxiliary camera, and transform the initial images taken by the auxiliary camera to the pixel plane of the main camera according to the pixel plane transformation matrix, so as to obtain a plurality of images to be stitched under the pixel plane of the main camera; A judging unit, used for projecting the fields of view of all cameras into the world coordinate system and judging whether the projection positions of the fields of view overlap; an extraction unit, configured to construct an image pyramid based on the image to be stitched when the judgment result of the judgment unit is yes, extract feature points on each layer of the image pyramid by using a SURF feature extraction algorithm to obtain a feature point set; and filter the feature point set according to a distance threshold and a quality threshold to obtain a feature point pair set, wherein the distance threshold is used to filter and remove feature point pairs whose distance is less than the distance threshold, and the quality threshold is used to filter and remove feature points whose response values are less than the quality threshold; The stitching unit is used to calculate the perspective transformation matrix between the images to be stitched based on the set of feature point pairs, and to stitch and fuse the images to be stitched according to the perspective transformation matrix and the view projection position.
7. A device for splicing images taken by multiple cameras for appearance inspection, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, and when the program is executed on a computer, the method according to any one of claims 1 to 5 is performed.
Citation Information
Patent Citations
Method for screening matching pairs of feature points to splice images
CN102819835A
Panoramic video rapid splicing method and system
CN110782394A