High-speed scanning splicing method, system and device based on dynamic deformation compensation
By constructing a checkerboard calibration board that integrates QR code position encoding and using an inverse projection algorithm to expand the matching point pairs, combined with Deloni triangulation and a convolutional neural network optical flow estimation model, the dynamic deformation error and local misalignment accumulation problem of image stitching under high-speed scanning are solved, achieving efficient image stitching and meeting the high-precision requirements of printed circuit board inspection.
Patent Information
- Application Number
- CN202511172793.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing technologies for printed circuit board inspection suffer from dynamic deformation errors, local misalignment accumulation, and real-time bottlenecks, resulting in insufficient image stitching accuracy and efficiency, especially failing to meet high-precision inspection requirements under high-speed scanning.
A high-speed scanning and stitching method based on dynamic deformation compensation is adopted. By constructing a checkerboard calibration board that integrates QR code position encoding, the matching point pairs are expanded using the inverse projection algorithm, and the Deloni triangulation and convolutional neural network optical flow estimation model are combined to perform deformation compensation and pixel merging, thereby achieving high-precision image stitching.
It significantly improves the accuracy and efficiency of image stitching, especially under high-speed scanning, it can achieve sub-pixel level accuracy stitching, meeting the needs of automatic optical inspection of printed circuit boards, and increasing the stitching speed by more than 40%.
Smart Images

Figure CN120725862B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial visual inspection technology, and in particular to a high-speed scanning and stitching method, system and equipment based on dynamic deformation compensation. Background Technology
[0002] With the increasing precision of electronic devices, the demand for high-precision defect detection in printed circuit boards (PCBs) continues to rise. Automated Optical Inspection (AOI) systems need to achieve an imaging resolution of over 100 pixels per millimeter. Since PCBs are typically tens of centimeters in size, the required image length and width reach tens of thousands of pixels, making single-shot imaging impossible for conventional imaging systems. Therefore, a two-dimensional platform with a camera is needed for multi-area scanning, followed by image stitching technology to reconstruct a complete, high-precision image. However, existing technologies face the following core challenges:
[0003] (1) Dynamic deformation error: When the two-dimensional platform moves at high speed, the camera fixture undergoes non-rigid deformation due to acceleration and deceleration, and different PCB sizes correspond to different scanning paths, causing the traditional fixed calibration parameters to fail.
[0004] (2) Accumulation of local misalignment: The splicing method based on feature point matching is prone to matching drift in low texture areas (such as blank substrates). Under high speed, small misalignment will be amplified exponentially with the number of splicing times.
[0005] (3) Real-time bottleneck: Traditional optical flow methods (such as Lucas-Kanade) have high computational complexity and are difficult to meet the millisecond-level response requirements of AOI systems, while splicing algorithms based on fixed templates cannot adapt to dynamic path planning.
[0006] Current mainstream solutions mostly use a chessboard calibration board combined with an affine transformation splicing framework, but they have limitations in the following aspects:
[0007] (11) The calibration board QR code and checkerboard features only provide sparse matching points, which are difficult to describe the local nonlinear distortion caused by dynamic deformation;
[0008] (12) The servo motor parameters (such as acceleration and position feedback) are not linked with the image stitching algorithm, so deformation pre-compensation cannot be achieved;
[0009] (13) Pixel fusion in overlapping areas relies on simple weighted averaging, which does not eliminate splicing artifacts caused by affine transformation residuals.
[0010] Therefore, there is an urgent need for an efficient image stitching method that integrates dynamic hardware parameters and has the ability to adaptively compensate for local deformation, in order to support the high-precision detection requirements of AOI systems under high-speed scanning. Summary of the Invention
[0011] Based on the technical problems existing in the background technology, this invention proposes a high-speed scanning stitching method, system and equipment based on dynamic deformation compensation, which solves the problem of adaptive optimization of stitching parameters caused by error sources such as mechanical deformation caused by high-speed movement of two-dimensional platform and camera tilt angle, and significantly improves the accuracy and efficiency of image stitching.
[0012] The high-speed scanning and stitching method based on dynamic deformation compensation proposed in this invention includes:
[0013] A checkerboard calibration board integrating QR code position encoding is constructed. By recognizing and decoding the QR code, a matching point pair between the steady-state sub-image and the precise point position of the theoretical image coordinate system is constructed. Then, an extended matching point pair is obtained by using the inverse projection algorithm. The steady-state sub-image is a standard image captured when the servo motor controls the camera to move to the predetermined shooting position and enters a static stable state.
[0014] The trained deformation parameter prediction model is used to estimate the estimated tilt angle of the camera during high-speed scanning, so as to correct the imaging error of the dynamic sub-image during shooting. The dynamic sub-image is the deformed image obtained by the camera moving on a two-dimensional platform.
[0015] Using the extended matching point pairs as the vertex set for the input of the Deloni triangulation, the corrected dynamic sub-image is split into H triangular pieces, which are then mapped to the precise points in the theoretical image coordinate system to obtain the corrected dynamic sub-image, where H is an integer.
[0016] The overlapping regions of the corrected dynamic sub-images are merged using an optical flow estimation model based on a convolutional neural network to obtain the final scanned and stitched image.
[0017] Furthermore, the construction of the checkerboard calibration board integrating QR code position encoding, and the construction of matching point pairs between the steady-state sub-image and the theoretically precise point positions in the image coordinate system through QR code recognition and decoding, specifically involves:
[0018] Construct a black and white checkerboard calibration board, and place QR codes containing set position information in the area of the white squares. The set position information includes the checkerboard size, the row and column number of the QR code, and the edge distance of the QR code.
[0019] The specified location information is populated using a Key-Value data storage model;
[0020] The location of the QR code is obtained from the steady-state sub-image and decoded. The corresponding position of the white square is obtained from the string carried by the QR code. The position coordinate information of the black and white square corner points is derived, so as to obtain the matching point pair between the position of the steady-state sub-image and the precise position of the theoretical image coordinate system.
[0021] Furthermore, the extended matching point pairs are obtained using the inverse projection algorithm, specifically:
[0022] Extract the tic-tac-toe straight lines from the steady-state sub-image and solve for the intersection points of the tic-tac-toe straight lines with the boundary of the steady-state sub-image. and the four corner points of the edge of the steady-state sub-image ;
[0023] The overall affine matrix is solved by matching point pairs between the steady-state sub-image obtained from QR code recognition and decoding and the theoretically precise points in the image coordinate system. The intersection point is calculated in reverse. and the four corner points of the edge The corresponding points in the precise location of the theoretical image coordinate system are used to obtain new point pairs, which are then incorporated into the matching point pairs to form expanded matching point pairs.
[0024] Furthermore, using the expanded matching point pairs as the input vertex set for the Deloni triangulation, the corrected dynamic sub-image is divided into H triangular pieces, which are then mapped to their theoretically precise points in the image coordinate system to obtain the corrected dynamic sub-image, specifically:
[0025] The intersections of the checkerboard pattern in the steady-state subimage are used as the initial feature point set. , intersection and the four corner points of the edge Combination as a set of image boundary points Set of image boundary points The points in the loop are sorted according to their spatial topological relationships to construct a polygonal closed loop.
[0026] Constructing a vertex set Triangular meshes conforming to the Deloni condition are gradually constructed using an incremental approach;
[0027] Perform edge flipping optimization on narrow triangular pieces with an aspect ratio greater than 3:1;
[0028] Calculate the local affine transformation matrix for each triangle, and map the corrected dynamic sub-image texture to the precise points in the theoretical image coordinate system according to the triangle blocks, thereby obtaining the corrected dynamic sub-image.
[0029] Furthermore, the overlapping regions of the corrected dynamic sub-images are merged using an optical flow estimation model based on a convolutional neural network to obtain the final stitched image, specifically as follows:
[0030] The overlapping region is input into the optical flow estimation model to generate the initial optical flow field;
[0031] The initial optical flow field is used to compensate for the deformation of the overlapping region to generate an aligned image.
[0032] An adaptive weighted fusion algorithm is used to merge pixels in overlapping areas to obtain the final scanned and stitched image.
[0033] Furthermore, deformation compensation is performed on the overlapping region using the initial optical flow field to generate an aligned image, specifically as follows:
[0034] Construct a multi-scale image pyramid to optimize feature matching step by step;
[0035] Solve for the least-squares affine matrix while keeping the extended matching point pairs unchanged;
[0036] Overlapping regions are mapped using a least-squares affine matrix to generate aligned images.
[0037] Furthermore, the adaptive weighted fusion algorithm is used to merge overlapping region pixels, wherein the weight function of the adaptive weighted fusion algorithm is... Specifically:
[0038] ;
[0039] in, These are the x and y coordinates of the overlapping region in the theoretical image coordinate system. The nearest matching point, This is the smoothing coefficient.
[0040] Furthermore, the training process of the deformation parameter prediction model is as follows:
[0041] Obtain steady-state and dynamic sub-images, and construct matching point pairs for steady-state and dynamic sub-images using QR code recognition and decoding respectively;
[0042] Second-order kinematic differential equations for the camera and the two-dimensional platform are established to characterize the camera's deformation tilt angle and the acceleration of the two-dimensional platform.
[0043] Model the camera and the two-dimensional platform, use finite element simulation to generate camera deformation tilt angles under different accelerations, construct a deformation dataset, and use it for pre-training of the deformation parameter prediction model;
[0044] Based on the deformation parameter prediction model pre-trained using simulation data, different velocities, accelerations, and paths are used to approximate each image acquisition point, and dynamic sub-image sequences are captured. We constructed a fine-tuning dataset and fine-tuned the deformation parameter prediction model.
[0045] The trainable parameters of the deformation parameter prediction model are updated using the least squares error between matching point pairs.
[0046] Furthermore, the processor executes the computer program to implement the method described above.
[0047] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described above.
[0048] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.
[0049] The advantages of the high-speed scanning stitching method, system, and equipment based on dynamic deformation compensation provided by this invention are: it is suitable for stitching images obtained at high speeds, and is particularly applicable to high-speed scanning imaging and defect detection in the context of automated optical inspection (AOI) of printed circuit boards. By employing methods such as steady-state sub-image calibration, expanding matching point pairs using inverse projection algorithms, introducing Delaunay triangulation to refine the stitching region, and constructing second-order kinematic differential equations for modeling, the invention solves the problem of adaptive optimization of stitching parameters caused by error sources such as mechanical deformation due to high-speed motion of the two-dimensional platform and camera tilt angle, significantly improving the accuracy and efficiency of image stitching. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the process of the present invention;
[0051] Figure 2 To assemble the flowchart at runtime;
[0052] Figure 3 To expand the flowchart for constructing matching point pairs, where Figure 3 The triangles in the notes represent coordinate anchor points, which correspond to the precise points described in this embodiment;
[0053] Figure 4 This is a flowchart of the training process for the deformation parameter prediction model. Detailed Implementation
[0054] The technical solution of the present invention will now be described in detail through specific embodiments. Many specific details are set forth in the following description to provide a thorough understanding of the invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0055] like Figures 1 to 4As shown, the high-speed scanning and stitching method based on dynamic deformation compensation proposed in this invention includes the following steps:
[0056] Step 1: Construct a checkerboard calibration board that integrates QR code position encoding. By recognizing and decoding the QR code, construct a matching point pair between the steady-state sub-image and the precise point position in the theoretical image coordinate system. Then, use the inverse projection algorithm to obtain an expanded matching point pair. The steady-state sub-image is the standard image (i.e., steady-state image) captured when the servo motor controls the camera to move to the predetermined shooting point and enters a static stable state.
[0057] Step 2: Use the trained deformation parameter prediction model to estimate the estimated tilt angle of the camera during high-speed scanning, so as to correct the imaging error of the dynamic sub-image during shooting. The dynamic sub-image is the deformed image (i.e., the dynamic image containing the deformation error of the actuator) obtained by the camera moving on the two-dimensional platform.
[0058] Step 3: Using the expanded matching point pairs as the vertex set of the Deloni triangulation input, the corrected dynamic sub-image is split into H triangular pieces, and each is mapped to the precise point position of the theoretical image coordinate system to obtain the corrected dynamic sub-image, where H is an integer;
[0059] Step 4: Use an optical flow estimation model based on a convolutional neural network to merge the pixels in the overlapping areas of the corrected dynamic sub-image to obtain the final scanned and stitched image.
[0060] This embodiment uses a trained deformation parameter prediction model to predict the camera tilt angle and the non-rigid deformation of the 2D platform in real time. Combined with servo motor parameter correction affine transformation, it solves the cumulative error caused by high-speed motion. The boundary is expanded to obtain expanded matching point pairs through the inverse projection algorithm, ensuring high coverage matching point pairs of Delaunay triangulation. This embodiment achieves a stitching speed improvement of more than 40% at sub-pixel accuracy through the above scanning stitching method, while supporting variable path planning.
[0061] In one embodiment, step one is to construct a checkerboard calibration board that integrates QR code position encoding, and construct matching point pairs between the steady-state sub-image and the precise points in the theoretical image coordinate system through QR code recognition and decoding, thereby using the inverse projection algorithm to obtain expanded matching point pairs.
[0062] Among them, a checkerboard calibration board integrating QR code position encoding is constructed, and a matching point pair between the steady-state sub-image and the precise point position of the theoretical image coordinate system is constructed through QR code recognition and decoding, specifically including (a1) to (a3):
[0063] (a1) Construct a black and white checkerboard calibration board, and place a QR code containing set position information in the area of the white square. The set position information includes the checkerboard size, the row and column number of the QR code, and the edge distance information of the QR code.
[0064] (a2) The set location information is populated using a Key-Value data storage model;
[0065] The key is used to uniquely identify the data, and the value is the corresponding specific content. At the same time, the checkerboard size, the row and column number of the QR code, and the margin information of the QR code are structured and encoded in JSON format. The key is sorted in ascending order of ASCII code to ensure the uniqueness of the QR code encoding.
[0066] (a3) Obtain the position of the QR code from the steady-state sub-image and decode it. Obtain the position of the white square from the string carried by the QR code. Derive the position coordinate information of the black and white square corner points. Based on this, obtain the matching point pair between the position coordinate information of the steady-state sub-image and the precise position of the theoretical image coordinate system. ,in The location points of the steady-state sub-image. These are the precise points in the theoretical image coordinate system (i.e., the coordinate anchor points in the theoretical image coordinate system).
[0067] In this embodiment, The preferred coordinates of the black and white grid corner points are used as the positions of the steady-state sub-image, but it is not ruled out that additional calibration points may be set on the checkerboard calibration board as the positions of the steady-state sub-image; furthermore, when the checkerboard calibration board is generated, the actual physical positions of each precise point on the calibration board are known, i.e. The checkerboard markings were known when the checkerboard was generated, therefore arrive The conversion only requires one PPI (Pixels Per Inch).
[0068] It should be noted that the process of acquiring steady-state sub-images is as follows: the camera is controlled by a servo motor to move at a low speed to the predetermined shooting position. After the mechanical structure stabilizes (i.e., it enters a static stable state), the camera is triggered to take a picture via the Modbus bus, thereby obtaining a steady-state sub-image. The above process is repeated to obtain steady-state sub-images corresponding to different predetermined shooting positions, thereby obtaining a sequence of steady-state sub-images for all predetermined shooting positions.
[0069] As an example of steady-state sub-image acquisition: a servo motor is used to control the camera to move slowly to the shooting point, and the system and mechanical structure are allowed to stabilize. During shooting, the speed should be guaranteed. acceleration And stable time By controlling the camera via a Modbus bus to take pictures, a standard steady-state sub-image of a single point can be obtained. Repeating the above steps several times will yield a sequence of standard steady-state sub-images for all shooting points.
[0070] In addition, the process of acquiring dynamic sub-images is as follows: by controlling the camera to move at high speed through a servo motor to perform high-speed continuous scanning, dynamic sub-images under different triggers can be obtained.
[0071] Therefore, the main differences between the steady-state sub-image and the dynamic sub-image are shown in Table 1:
[0072] Table 1
[0073]
[0074] During 2D decoding, the corresponding positions of the white squares are obtained from the JSON string carried by the QR code, thereby deriving the position information of the black and white square corners. The black and white squares refer to the black and white squares in the chessboard (i.e., the squares with QR codes), and the black and white square corners refer to the connection points between the black and white squares.
[0075] In addition, the extended matching point pairs, including (b1) to (b2), are obtained using the inverse projection algorithm:
[0076] (b1) Extract the tic-tac-toe straight lines of the checkerboard pattern from the steady-state sub-image. Solve for the intersection points of the grid-shaped straight lines and the boundary of the steady-state sub-image. and the four corner points of the edge of the steady-state sub-image ,in, All are symbolic markers, serving as point indices;
[0077] Among them, the four corner points of the edge of the steady-state sub-image Specifically, the four edge points of the steady-state sub-image are the upper left, upper right, lower right, and lower left, thus satisfying the polygon boundary closure condition of the Delaunay triangulation.
[0078] (b2) Matching point pairs between the steady-state sub-image obtained by QR code recognition and decoding and the theoretically precise point positions in the image coordinate system. Solving the global affine matrix Calculate the intersection point in reverse and the four corner points of the edge The corresponding points in the theoretical image coordinate system and Based on this, new point pairs are obtained, and these new point pairs are merged into the matching point pairs to form expanded matching point pairs. .
[0079] Each expression should satisfy:
[0080] ;
[0081] ;
[0082] ;
[0083] ;
[0084] ;
[0085] in These are the transformed target coordinates. These are the source coordinates before the transformation. The first straight line in the grid pattern A straight line in the row, The first straight line in the grid pattern The straight line of the column, These represent the total number of rows and columns of the grid pattern, respectively. These represent the width and length of the steady-state sub-image, respectively. Points in the steady-state sub-image Using the global affine matrix A point mapped in the theoretical image coordinate system. Specifically or , For boundary operators, This indicates the boundary region of the image.
[0086] In one embodiment, step two, estimating the estimated tilt angle of the camera during high-speed scanning using the trained deformation parameter prediction model, and thereby correcting the imaging error of the dynamic sub-image during capture, specifically involves:
[0087] The training process for the deformation parameter prediction model in this embodiment is as follows:
[0088] (c1) Obtain the steady-state sub-image and the dynamic sub-image, and construct matching point pairs for the steady-state sub-image and the dynamic sub-image respectively using QR code recognition and decoding;
[0089] It should be noted that the steady-state and dynamic sub-images used for training the deformation parameter prediction model are data from the fine-tuning dataset. The steady-state and dynamic sub-images used in this training are sourced in the same way as those used in the high-speed scanning process, namely, steady-state sub-images under low-speed steady-state conditions and dynamic sub-images obtained at different speeds.
[0090] (c2) Establish the second-order kinematic differential equations for the camera and the two-dimensional platform to characterize the camera's deformation tilt angle and the acceleration of the two-dimensional platform. The second-order kinematic differential equations are as follows:
[0091] ;
[0092] in, For rotational inertia, This is the stiffness coefficient. The damping coefficient is... The torque coefficient, To accelerate the real-time platform, The time after the start of the exercise. It is the estimated tilt angle of the camera. , These are the trainable parameters for the deformation parameter prediction model.
[0093] (c3) Model the camera and the two-dimensional platform, generate the camera deformation tilt angle under different accelerations using finite element simulation, construct a deformation dataset based on it, and use it for pre-training of the deformation parameter prediction model;
[0094] The camera and 2D platform were modeled using 3D modeling software (such as SolidWorks), and multiple accelerations were calculated using the deformation dataset generated by the finite element simulation of the 3D modeling software. and time Estimated tilt angle of combined cameras This generates the dynamic deformation parameters required during the pre-training of the deformation parameter prediction model. Thus, a pre-trained deformation parameter prediction model is obtained, in which, This indicates the camera's tilt angle.
[0095] It should be noted that, since it is difficult to collect actual speed, acceleration and path data in large quantities as in simulation, this embodiment first uses simulation data to pre-train the deformation parameter prediction model.
[0096] (c4) Based on the deformation parameter prediction model pre-trained using simulation data, different velocities, accelerations, and paths are used to approximate each image acquisition point, and dynamic sub-image sequences are captured. We constructed a fine-tuning dataset and fine-tuned the deformation parameter prediction model.
[0097] For a specific two-dimensional image acquisition system, different servo motor motion parameters (i.e., speed, acceleration, and path, etc.) are used to approximate each image acquisition point. That is, the matching point pairs of steady-state sub-images and dynamic sub-images, as well as the servo motor motion parameters, are used as fine-tuning datasets to fine-tune the pre-trained deformation parameter prediction model. The deformation of the camera-two-dimensional platform during high-speed motion is calculated, thereby correcting the imaging error caused by the deformation of the actuator during dynamic sub-image capture.
[0098] (c5) Update the trainable parameters of the deformation parameter prediction model using the least squares error between matching point pairs.
[0099] It should be noted that when using the trained deformation parameter prediction model to estimate the camera's tilt angle, the motion parameters of the servo motor (speed, path, acceleration, etc.) during actual data acquisition are used as input to the pre-trained deformation parameter prediction model to assist in generating the estimated tilt angle of the camera.
[0100] In one embodiment, step three involves using the expanded matching point pairs as the vertex set for the Deloni triangulation input, splitting the corrected dynamic sub-image into H triangular patches, and mapping each patch to its theoretically precise point in the image coordinate system to obtain the corrected dynamic sub-image. Specifically:
[0101] (d1) Use the intersections of the checkerboard pattern in the steady-state sub-image as the initial feature point set. , intersection and the four corner points of the edge Combination as a set of image boundary points Set of image boundary points The points in the diagram are sorted according to their spatial topological relationships to construct a polygonal closed loop. ;
[0102] (d2) Constructing the vertex set Triangular meshes conforming to the Deloni condition are gradually constructed using an incremental approach;
[0103] Initial feature point set and image boundary point set Merging yields a vertex set The Bowyer-Watson algorithm is used to perform incremental triangulation on the corrected dynamic sub-image, satisfying the empty circumcircle criterion.
[0104] Among them, the Bowyer-Watson algorithm is a point-by-point insertion algorithm for generating Delaunay triangulations. Its core idea is to incrementally construct triangular meshes that meet the Delaunay conditions.
[0105] (d3) Perform edge flipping optimization on narrow triangular pieces with an aspect ratio greater than 3:1:
[0106] ;
[0107] in, These are the three vertices of the triangle. These are the three vertices after performing a pass-through optimization on the triangle. These are the three side lengths of the triangle.
[0108] Virtual points are inserted in sparse texture regions (where a single triangle covers more than one-tenth of the area of the corrected dynamic sub-image). The virtual points are selected as the centroids of the triangles, and the corresponding final stitched image is calculated using bilinear interpolation. Corresponding coordinates:
[0109] ;
[0110] in, For all triangles generated by triangulation , For the area of the corresponding triangle, The total area of the currently processed sub-image. The coordinates of the virtual vertex to be inserted are determined by selecting the centroid of the triangle. Triangles The coordinates of the three vertices.
[0111] The meaning of the above mathematical expression is: for any segmented triangle When its area is greater than one-tenth of the area of the corrected dynamic sub-image, an additional virtual point is added. This allows for further subdivision of the triangle (using Delaunay triangulation).
[0112] (d4) Calculate the local affine transformation matrix for each triangle, and map the corrected dynamic sub-image texture to the precise point of the theoretical image coordinate system according to the triangle blocks, thereby obtaining the corrected dynamic sub-image.
[0113] This embodiment uses (d1) to (d4) to divide the corrected dynamic sub-image into several triangular sub-regions and perform affine mapping on each, which can largely solve the nonlinear imaging error caused by the nonlinear distortion of camera imaging and the rolling shutter effect during dynamic shooting.
[0114] In one embodiment, step four involves using a convolutional neural network-based optical flow estimation model to merge pixels in the overlapping regions of the corrected dynamic sub-image, thereby obtaining the final scanned and stitched image. Specifically:
[0115] (e1) Input the overlapping region into the optical flow estimation model to generate the initial optical flow field;
[0116] The overlapping region (denoted as image) and images The offset is converted into a dense displacement field, which is then used as the initial optical flow estimate input and appended to the optical flow estimation model (PWC-Net) to generate the initial optical flow field. Its optical flow optimization is denoted as ;
[0117] ;
[0118] in, For the current estimated optical flow field to be optimized, In the first The original input image at each level of the image pyramid. These are images of two overlapping regions. For feature extraction function, These are the regularization weight coefficients. For the initial optical flow field, It is an L2 norm.
[0119] The above The formula describes an optical flow optimization process. Its goal is to find an optimal dense displacement field (optical flow field) that satisfies two conditions simultaneously:
[0120] Feature matching consistency: Image Based on current estimated optical flow To deform (i.e.) After that, the extracted features Should be related to the image Features As similar as possible;
[0121] Minimize the deviation from the initial estimate: find the displacement field It should not deviate from the given initial optical flow estimate Too far.
[0122] (e2) The initial optical flow field is used to perform deformation compensation on the overlapping area to generate an aligned image;
[0123] In the above optical flow optimization During the process, multi-scale image pyramids were constructed. The feature matching is optimized step by step. Under the constraint that the expanded matching point pairs remain unchanged, the least squares affine matrix is solved. The overlapping regions are mapped using the least squares affine matrix to generate an aligned image.
[0124] Multiscale image pyramid for:
[0125] ;
[0126] in, The image is input into the optical flow estimation model. The index is the index of the input image. In this embodiment, it refers to the image index of the overlapping area, i.e., 1 and 2. However, this embodiment does not exclude the possibility that there are more than two images in the overlapping area. For image pyramid hierarchical indexing, This represents the top layer (the coarsest and lowest resolution scale). Indicates the intermediate layer. This represents the lowest level (the finest and highest resolution scale, usually close to or equal to the original image resolution). The larger it is, the finer the scale. For downsampling operation, For image In the image pyramid The image corresponding to the layer is an image. Low-resolution representation at this scale.
[0127] This embodiment sets up an image pyramid structure from coarse to fine, at the coarse scale... The image resolution is low, but the features are large, making it easy to find approximate matches and resulting in high computational efficiency. The solved least-squares affine matrix captures a large range of global motion.
[0128] This embodiment constructs multi-scale image pyramid structures with resolutions of 1 / 2, 1 / 4, and 1 / 8, and optimizes optical flow step by step from low resolution to find feature matching point pairs for overlapping regions.
[0129] The formula for solving the least squares affine matrix is:
[0130] ;
[0131] in, Here is the homogeneous representation of the least-squares affine matrix to be solved. Given the constraints, the solution matrix is required to be obtained. rank It must be equal to 3.
[0132] It should be noted that in optical flow optimization During the process, the optical flow field is optimized through one level of the image pyramid. This will serve as the initial optical flow field for the next, more refined pyramid level. Optimize the optical flow field step by step At each scale, find the overlapping regions (images). and images We find the matching feature point pairs between the points, and under the condition that the set of expanded matching point pairs remains unchanged at the current scale, we solve for an optimal affine transformation matrix that best aligns these expanded matching point pairs in the least squares sense. We then use the solved affine matrix... For images in overlapping regions Perform transformations to generate an image similar to the reference image. Aligned new image; optical flow field optimized for the current scale As an initial value, it is passed to the next level of more refined scale for further optimization.
[0133] (e3) An adaptive weighted fusion algorithm is used to merge pixels in overlapping areas to obtain the final scanned and stitched image.
[0134] The weight function of the adaptive weighted fusion algorithm mentioned above Specifically:
[0135] ;
[0136] in, These are the x and y coordinates of the overlapping region in the theoretical image coordinate system. The nearest matching point, This is the smoothing coefficient.
[0137] An adaptive weighted fusion algorithm is used to merge pixels in overlapping regions to obtain a fused image of the overlapping regions. :
[0138] ;
[0139] in, and respectively through affine matrices For images and images The image obtained after performing an affine transformation. and The weights corresponding to the pixels in the overlapping region are respectively
[0140] To verify the performance of this embodiment, multi-size samples were collected from a PCB manufacturing plant's production line for testing. The test conditions are shown in Table 2.
[0141] Table 2
[0142]
[0143] Table 3 shows the comparison results between this embodiment and the traditional method (stitching error unit: pixels, time unit: ms / sub-image):
[0144] Table 3
[0145]
[0146] Traditional affine transformation involves finding three matching points on the image, estimating the affine matrix, and then stitching them together. This is the single affine transformation method described in the paper "Design of Affine Transformation Algorithm Based on Triangulation Network". This method is the most basic and simplest image stitching method, but it has the worst performance. SIFT feature matching is based on the SIFT algorithm to extract feature values, then filters the feature points for matching points, and finally estimates the affine matrix for stitching. It is also a basic method, as described in the paper "Automatic panoramic image stitching using invariant features". Optical flow method mainly achieves registration by estimating the motion vectors of pixels between consecutive frames. It first calculates the optical flow field and then performs motion estimation (but it cannot consider deformation due to high-speed motion). For example, the method disclosed in the paper "Video Stitching Based on Optical Flow" is used. However, this method cannot utilize prior information such as the motion parameters of the two-dimensional platform and deformation estimation, and its stitching accuracy is lower than that of this embodiment.
[0147] Experiments show that this embodiment improves the stitching efficiency by up to 67.1% under the premise of sub-pixel accuracy (≤1.0 pixel), meeting the requirements of high-speed automatic optical inspection (AOI).
[0148] Therefore, this embodiment is used for stitching images obtained at high speeds, and is particularly suitable for high-speed scanning imaging and defect detection in automated optical inspection (AOI) scenarios for printed circuit boards. This embodiment solves the problem of adaptive optimization of stitching parameters caused by error sources such as mechanical deformation and camera tilt angle due to high-speed motion of the two-dimensional platform by methods such as steady-state sub-image calibration, expanding matching point pairs using inverse projection algorithm, introducing Delaunay triangulation to refine the stitching region, and constructing second-order kinematic differential equation modeling. This significantly improves the accuracy and efficiency of image stitching.
[0149] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A high-speed scanning and stitching method based on dynamic deformation compensation, characterized in that, include: A checkerboard calibration board integrating QR code position encoding is constructed. By recognizing and decoding the QR code, a matching point pair between the steady-state sub-image and the precise point position of the theoretical image coordinate system is constructed. Then, an extended matching point pair is obtained by using the inverse projection algorithm. The steady-state sub-image is a standard image captured when the servo motor controls the camera to move to the predetermined shooting position and enters a static stable state. The trained deformation parameter prediction model is used to estimate the estimated tilt angle of the camera during high-speed scanning, so as to correct the imaging error of the dynamic sub-image during shooting. The dynamic sub-image is the deformed image obtained by the camera moving on a two-dimensional platform. Using the extended matching point pairs as the vertex set for the input of the Deloni triangulation, the corrected dynamic sub-image is split into H triangular pieces, which are then mapped to the precise points in the theoretical image coordinate system to obtain the corrected dynamic sub-image, where H is an integer. The overlapping regions of the corrected dynamic sub-images are merged using an optical flow estimation model based on a convolutional neural network to obtain the final scanned and stitched image.
2. The high-speed scanning and stitching method according to claim 1, characterized in that, The construction of the checkerboard calibration board integrating QR code position encoding involves using QR code recognition and decoding to construct matching point pairs between the steady-state sub-image and the theoretically precise point positions in the image coordinate system. Specifically: Construct a black and white checkerboard calibration board, and place QR codes containing set position information in the area of the white squares. The set position information includes the checkerboard size, the row and column number of the QR code, and the edge distance of the QR code. The specified location information is populated using a Key-Value data storage model; The location of the QR code is obtained from the steady-state sub-image and decoded. The corresponding position of the white square is obtained from the string carried by the QR code. The position coordinate information of the black and white square corner points is derived, so as to obtain the matching point pair between the position of the steady-state sub-image and the precise position of the theoretical image coordinate system.
3. The high-speed scanning and stitching method according to claim 1, characterized in that, The extended matching point pairs are obtained using the inverse projection algorithm, specifically as follows: Extract the tic-tac-toe straight lines from the steady-state sub-image and solve for the intersection points of the tic-tac-toe straight lines with the boundary of the steady-state sub-image. and the four corner points of the edge of the steady-state sub-image ; The original affine matrix is solved by matching point pairs between the steady-state sub-image obtained from QR code recognition and decoding and the theoretically precise points in the image coordinate system. The intersection point is calculated in reverse. and the four corner points of the edge Theoretically, the corresponding points of the precise points in the image coordinate system are used to obtain new point pairs, and the new point pairs are merged into the matching point pairs to form expanded matching point pairs.
4. The high-speed scanning and stitching method according to claim 3, characterized in that, Using the extended matching point pairs as the input vertex set for Deloni triangulation, the corrected dynamic sub-image is divided into H triangular patches, which are then mapped to their theoretically precise points in the image coordinate system to obtain the corrected dynamic sub-image, specifically: The intersections of the checkerboard pattern in the steady-state subimage are used as the initial feature point set. , intersection and the four corner points of the edge Combination as a set of image boundary points Set of image boundary points The points in the loop are sorted according to their spatial topological relationships to construct a polygonal closed loop. Constructing a vertex set Triangular meshes conforming to the Deloni condition are gradually constructed using an incremental approach; Perform edge flipping optimization on narrow triangular pieces with an aspect ratio greater than 3:1; Calculate the local affine transformation matrix for each triangle, and map the corrected dynamic sub-image texture to the precise points in the theoretical image coordinate system according to the triangle blocks, thereby obtaining the corrected dynamic sub-image.
5. The high-speed scanning and stitching method according to claim 3, characterized in that, The overlapping regions of the corrected dynamic sub-images are merged using an optical flow estimation model based on a convolutional neural network to obtain the final stitched image. Specifically: The overlapping region is input into the optical flow estimation model to generate the initial optical flow field; The initial optical flow field is used to compensate for the deformation of the overlapping region to generate an aligned image; An adaptive weighted fusion algorithm is used to merge pixels in overlapping areas to obtain the final scanned and stitched image.
6. The high-speed scanning and stitching method according to claim 5, characterized in that, Deformation compensation of the overlapping region is performed using the initial optical flow field to generate an aligned image, specifically as follows: Construct a multi-scale image pyramid to optimize feature matching step by step; Solve for the least-squares affine matrix while keeping the extended matching point pairs unchanged; Overlapping regions are mapped using a least-squares affine matrix to generate aligned images.
7. The high-speed scanning and stitching method according to claim 5, characterized in that, The method employs an adaptive weighted fusion algorithm to merge overlapping pixels, wherein the weight function of the adaptive weighted fusion algorithm is... Specifically: in, These are the x and y coordinates of the overlapping region in the theoretical image coordinate system. The nearest matching point, This is the smoothing coefficient.
8. The high-speed scanning and stitching method according to claim 1, characterized in that, The training process of the deformation parameter prediction model is as follows: Obtain steady-state and dynamic sub-images, and construct matching point pairs for steady-state and dynamic sub-images using QR code recognition and decoding respectively; Second-order kinematic differential equations for the camera and the two-dimensional platform are established to characterize the camera's deformation tilt angle and the acceleration of the two-dimensional platform. Model the camera and the two-dimensional platform, use finite element simulation to generate camera deformation tilt angles under different accelerations, construct a deformation dataset, and use it for pre-training of the deformation parameter prediction model; Based on the deformation parameter prediction model pre-trained using simulation data, different velocities, accelerations, and paths are used to approximate each image acquisition point, and dynamic sub-image sequences are captured. We constructed a fine-tuning dataset and fine-tuned the deformation parameter prediction model. The trainable parameters of the deformation parameter prediction model are updated using the least squares error between matching point pairs.
9. A computer system comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-8.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-8.
Citation Information
Patent Citations
High-robustness and high-precision camera calibration board and angular point detection method
CN118674789A
Hand-eye calibration method, apparatus and device based on point cloud registration
WO2024207704A1