Silk Cake Defect Detection Method Based on Improved LoFTR Algorithm
Through the improved LoFTR algorithm and camera calibration technology, the problem of insufficient feature point extraction capability and poor splicing effect of image splicing in silk cake surface defect detection is solved, and efficient and accurate detection of silk cake surface defects is achieved.
Patent Information
- Application Number
- CN202311037105.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-08-17
AI Technical Summary
In the detection of surface defects of silk cakes, especially when multi-camera acquires images, there are problems such as insufficient feature point extraction capabilities, poor splicing effects, and inability to meet real-time requirements.
The improved LoFTR algorithm is used to calibrate and distortion correction through standard dot calibration plates, feature points are extracted and matched, and the homography matrix is solved using the PROSAC algorithm, perspective transformation and weighted average image fusion are performed to eliminate splicing seams.
The feature point extraction capability and splicing and fusion speed are improved, and the efficiency and accuracy of surface defect detection of silk cakes are achieved, meeting the requirements of real-time.
Smart Images

Figure CN116993709B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cake image detection, and particularly relates to a method for detecting cake defects based on an improved LoFTR algorithm. Background Art
[0002] A cake is formed by regularly winding filament yarns around a paper tube for convenient transportation, storage, and packaging. During the preparation of the cake, various defects may occur due to reasons such as too high winding speed of the machine and uneven tension. The existence of defects seriously affects the quality, appearance, and usability of the cake. Currently, there are two methods for detecting surface defects of the cake: manual detection and visual detection. Manual detection is inefficient, requires high skills from the inspectors, and is tiring. In addition, some large manufacturers use machine vision for detection. However, since the cake has multiple surfaces and each surface is relatively large, it is difficult to capture the entire surface using a single industrial camera. Therefore, a multi-camera synchronization method is used to capture the surfaces of the cake.
[0003] The grading standard of the surface quality of the cake is related to the type and quantity of surface defects. When using multiple cameras for acquisition, there will be overlapping regions in the images captured by each camera, resulting in repeated detection of defects in this part, which will affect the quality grading of the cake. Therefore, it is necessary to stitch and fuse the images captured by multiple cameras, eliminate the overlapping regions, and then transmit them to the upper computer for detection.
[0004] During the actual detection process of the cake, the proportion of defective cakes is relatively small. It takes a long time to stitch each captured image. Therefore, it is necessary to first determine whether the cake has defects and then consider whether to stitch its image, reducing the time consumption and enabling correct grading of the surface quality of the cake.
[0005] Currently, there is little research on the stitching method for such weakly textured images on the surface of the cake. Some traditional methods such as SIFT, SURF, and ORB algorithms have weak feature point extraction capabilities and slow extraction speeds for weakly textured surfaces. As a result, it is sometimes difficult to solve the optimal homography transformation matrix between two images, leading to poor image stitching effects and inability to meet the real-time requirements, and thus cannot be applied to the stitching of cake surface defect images.
[0006] The patent with the publication number CN116152073A discloses an improved multi-scale fundus image stitching method based on the Loftr algorithm. The captured fundus images are input into the original Loftr algorithm to complete feature point extraction and feature point matching, and then the homography matrix is solved using the obtained feature points to achieve image stitching. Although this method has stronger extraction capabilities compared to traditional feature point extraction algorithms, its extraction capabilities for some weakly textured images may still be insufficient, affecting the subsequent stitching effect.
[0007] The patent with publication number CN115526781A discloses a stitching method based on overlapping areas of images. The method uses the SURF algorithm to extract feature points and generate feature point descriptors from images, uses the FLANN algorithm to perform rough matching to obtain a rough matching point pair group, uses the distance screening method to screen the obtained rough matching point pair group to obtain a fine matching point pair group, uses the fine matching point pair group to establish a transformation matrix model using the RANSAC algorithm, and performs image stitching; uses the weighted fusion method to optimize the gaps in the stitched images, eliminate the stitching seams, and complete the image fusion. This method takes a long time to stitch images, and the SURF algorithm is not capable of extracting feature points from weak texture images. Summary of the invention
[0008] In view of the problems existing in the prior art, the purpose of the embodiments of the present application is to provide a silk cake defect detection method based on an improved LoFTR algorithm, which can complete the stitching of collected silk cake images with repeated areas in a relatively short time. Compared with the existing algorithms, it has great advantages in the ability to extract feature points and the speed of image stitching and fusion.
[0009] According to a first aspect of an embodiment of the present application, a method for detecting wire cake defects based on an improved LoFTR algorithm is provided, comprising:
[0010] (1) Calibrate each camera in turn based on the standard dot calibration plate to obtain the intrinsic parameter matrix, extrinsic parameter matrix and distortion coefficient of each camera, so as to perform distortion correction for each camera;
[0011] (2) obtaining multiple images of the surfaces of the silk cakes collected by multiple cameras after distortion correction, and obtaining a plurality of image pairs to be stitched with repeated areas;
[0012] (3) performing defect detection on the silk cake image. If there are hairy threads, hairiness, broken threads, or looped threads on the surface of the silk cake, the silk cake is defective. If there are no defects, the inspection is completed. If there are defects, proceed to the next step.
[0013] (4) pre-estimating the overlapping area of the pair of images to be stitched according to camera parameters to obtain the proportion of the overlapping area to the image;
[0014] (5) performing pixel adjustment and exposure compensation on the image pair to be stitched to facilitate stitching and fusion;
[0015] (6) According to the ratio of the overlapping area to the image, use masks of corresponding ratios to extract the overlapping areas of the image pairs to be stitched after processing in step (5), use the improved Transformer-based LoFTR algorithm to extract feature points from the overlapping areas, and remove incorrectly matched feature point pairs according to the confidence of the feature point pairs to obtain effectively matched feature point pairs;
[0016] (7) Based on the effectively matched feature point pairs, use the PROSAC algorithm for random sampling, and then solve the homography matrix of the image pairs to be stitched processed in step (5);
[0017] (8) Use the obtained homography matrix to perform perspective transformation on the right image to be stitched in the image pairs to be stitched processed in step (5), and copy the left image to be stitched onto the perspectively transformed right image to be stitched;
[0018] (9) Use the weighted average image fusion algorithm to fuse the overlapping regions of the directly stitched images obtained in step (8), eliminate the stitching seam, and obtain an image of the complete and non-repetitive region on the surface of the cake;
[0019] (10) Perform defect detection on the images of the complete and non-repetitive regions on each surface, determine the type, position, and number of defects, so as to grade the quality of the cake.
[0020] Further, the multi-camera acquisition device for acquiring cake images includes a cake support base, a circumferential surface acquisition camera, an upper end surface acquisition camera, and a lower end surface acquisition camera. There are two upper end surface acquisition cameras and two lower end surface acquisition cameras respectively, and there are 4 circumferential surface acquisition cameras.
[0021] Further, in step (4),
[0022] The overlapping regions between the cake end surface images satisfy:
[0023]
[0024] where fov is the field of view angle of the camera, l is the distance between the two cameras, h is the distance from the camera to the surface of the cake, and φ is the proportion of the overlapping region of the cake end surface image in the original image area;
[0025] The overlapping regions between the cake circumferential surface images satisfy:
[0026]
[0027] where λ is the proportion of the overlapping region of the cake circumferential surface image, fov is the field of view angle of the camera, and β is the angle between the two cameras.
[0028] Further, in step (5), the pixel adjustment is to use pixel points with a value of 0 to perform edge filling on the images acquired by multiple cameras;
[0029] The exposure compensation is performed by using gain exposure, and the gain coefficient is realized through the error function e:
[0030]
[0031] where
[0032]
[0033] g i and g j are the gain coefficients of images i and j, R(i, j) is the overlapping part of the two images, R, G, and B are the intensity values of each channel component of the image, N ij is the number of pixels in the overlapping part, σ N and σ g are the standard deviations of the error and the gain respectively. Usually, σ N = 10, σ g = 0.1, I ij is the average value of the image intensity in the overlapping part, n 1 and n 2 are the number of pixels in the width and height of the overlapping part respectively.
[0034] Furthermore, in step (6), the improved Transformer-based LoFTR algorithm is implemented by a local feature CNN extraction module, a feature transformation module, a coarse-grained matching module, and a fine-grained matching module, where:
[0035] The local feature CNN extraction module uses the ResNetFPN network to perform three downsamplings on the input two cake images to the levels of 1 / 2, 1 / 4, and 1 / 8 of the image size. The coarse-grained image feature map at the 1 / 8 level is upsampled twice using the bilinear interpolation method to obtain the feature maps at 1 / 4 and 1 / 2 levels in sequence, and is fused with the high-level feature information to output the coarse-grained image feature map at the 1 / 8 level and the fine-grained image feature map at 1 / 2 where an scSE attention mechanism is added before each downsampling;
[0036] The feature transformation module is used to extract the local features related to the position and context from the incoming coarse-grained image feature map and
[0037] The coarse-grained matching module uses a differentiable matching layer to match the features and obtain the confidence matrix P c , and uses the mutual nearest neighbor algorithm to select the matching pairs to obtain the coarse-grained matching prediction M c ;
[0038] For each coarse-grained matching pair, the fine-grained matching module extracts from the fine-grained image feature map in Extract a local window of size w×w at the position, input it into the self-attention and cross-attention layers, and use softmax to calculate the score matrix and matching probability matrix of the window. Finally, obtain the coordinates of the final feature matching points in the fine-grained image feature map and output.
[0039] Furthermore, in step (7), the homography matrix H is calculated by the following formula and the matrix that minimizes the back-projection error rate is selected:
[0040]
[0041] where (x i , y i ) are the coordinates of the feature points sampled in the left image to be stitched in the overlapping area, and (x′ i , y′ i ) are the coordinates of the feature points sampled in the right image to be stitched.
[0042] Furthermore, in step (8), for the coordinates (u, v) of the right image to be stitched that undergoes perspective transformation, the perspective transformation results in the image coordinates (x, y):
[0043]
[0044]
[0045] [x′, y′, w′] = [u, v, w]H
[0046] where H is the homography matrix.
[0047] Furthermore, in step (9), use a weighted average image fusion algorithm to fuse the overlapping area of the directly stitched images in step (8), specifically:
[0048]
[0049] where g(x, y) is the pixel of the fused image point, g 1 (x, y) is the pixel point of the left image to be stitched, g 2 (x, y) is the pixel point of the right image to be stitched, g 1 ∩g 2 is the overlapping part between the images to be stitched, and ω is the weight value.
[0050] According to the second aspect of the embodiments of the present application, there is provided an electronic device, including:
[0051] One or more processors;
[0052] A memory for storing one or more programs;
[0053] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.
[0054] According to a third aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in the first aspect are implemented.
[0055] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0056] As can be seen from the above embodiments, the present application proposes a method for detecting cake defects based on an improved LoFTR algorithm for the problem of stitching cake images collected by multiple cameras, and solves the problems of insufficient extraction ability of the existing keypoint extraction algorithm for weakly textured images, inability to collect the whole cake with a single camera due to its large size, repeated counting of defects caused by overlapping image regions collected by multiple cameras, and long required time, enhances the keypoint extraction ability, and improves the stitching and fusion speed.
[0057] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0059] Figure 1 It is a schematic flow chart of a method for seamless fusion of cake defect images collected by multiple cameras based on an improved LoFTR algorithm;
[0060] Figure 2 It is a schematic diagram of collecting images of each surface of a cake in an exemplary embodiment, where (a) is a side view of collecting images of each surface of the cake, and (b) is a top view of collecting images of each surface of the cake;
[0061] Figure 3 It is a cake image to be stitched shown in an exemplary embodiment, where (a) is an end face image of the cake, and (b) is a circumferential face image of the cake;
[0062] Figure 4 It is an overall framework diagram of the improved LoFTR algorithm;
[0063] Figure 5 It is a network structure diagram of the feature extraction module in the improved LoFTR algorithm;
[0064] Figure 6 It is a schematic diagram of the principle of the scSE attention mechanism;
[0065] Figure 7 It is a matching graph of feature point pairs extracted by improving the LoFTR algorithm;
[0066] Figure 8 It is a matching graph of feature point pairs after confidence filtering;
[0067] Figure 9 It is a direct splicing graph after the perspective transformation of the image pairs to be spliced. Among them, (a) is a schematic diagram of the direct splicing of the end face of the silk cake, and (b)-(d) are schematic diagrams of the direct splicing of the circumferential face of the silk cake;
[0068] Figure 10 It is a result graph after image fusion using the weighted average method. Among them, (a) is a schematic diagram of the image fusion of the end face of the silk cake, and (b)-(d) are schematic diagrams of the image fusion of the circumferential face of the silk cake;
[0069] Figure 11 It is a block diagram of a seamless fusion device for silk cake defect images collected by multiple cameras based on the improved LoFTR algorithm;
[0070] Figure 12 It is a schematic diagram of an electronic device.
[0071] Reference numerals: silk cake support base 1, silk cake 2, circumferential surface acquisition camera 3, upper end face acquisition camera 4, lower end face acquisition camera 5. Detailed implementation manners
[0072] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application.
[0073] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms of "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0074] It should be understood that although terms such as first, second, and third may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0075] Figure 1 is a flowchart of a method for detecting cake defects based on an improved LoFTR algorithm shown according to an exemplary embodiment, as Figure 1 shown, this method is applied to a terminal and may include the following steps:
[0076] (1) Calibrate each camera in turn based on a standard dot calibration board to obtain the internal parameter matrix, external parameter matrix, and distortion coefficient of each camera, so as to perform distortion correction on each camera;
[0077] (2) Obtain multiple cake images of each surface of the cake collected by multiple cameras after distortion correction to obtain a number of pairs of images to be stitched with overlapping regions;
[0078] (3) Perform defect detection on the cake images. If there are fuzz, flyers, broken wires, or looped wires on the cake surface, it is considered defective. If there is no defect, this detection is completed. If there is a defect, continue to the next step;
[0079] (4) Pre-estimate the overlapping region of the pairs of images to be stitched according to the camera parameters to obtain the proportion of the overlapping region in the image;
[0080] (5) Perform pixel adjustment and exposure compensation on the pairs of images to be stitched to facilitate stitching and fusion;
[0081] (6) According to the proportion of the overlapping region in the image, use corresponding proportion masks to extract the overlapping regions of the pairs of images to be stitched processed in step (5), use the improved Transformer-based LoFTR algorithm to extract feature points from the overlapping regions, and eliminate wrongly matched feature point pairs according to the confidence of the feature point pairs to obtain effectively matched feature point pairs;
[0082] (7) Based on the effectively matched feature point pairs, use the PROSAC algorithm for random sampling, and then solve the homography matrix of the pairs of images to be stitched processed in step (5);
[0083] (8) Use the obtained homography matrix to perform perspective transformation on the right image to be stitched in the pair of images to be stitched obtained in step (5), and copy the left image to be stitched onto the right image to be stitched after perspective transformation;
[0084] (9) Use the weighted average image fusion algorithm to fuse the overlapping regions of the directly stitched images obtained in step (8) to eliminate the stitching seam and obtain an image of the complete and non-repetitive region on the surface of the cake;
[0085] (10) Perform defect detection on the images of each complete and non-repetitive region on the surface to determine the type, position and number of defects, so as to grade the quality of the cake.
[0086] As can be seen from the above embodiments, the present invention proposes a method for detecting cake defects based on an improved LoFTR algorithm for the problem of stitching images of cakes collected by multiple cameras, which solves the problems of insufficient extraction ability of existing key point extraction algorithms for weakly textured images, incomplete collection of large-sized cakes by a single camera, repeated statistics of defects caused by overlapping image regions collected by multiple cameras, and long required time, enhances the feature point extraction ability, and improves the stitching and fusion speed.
[0087] In the specific implementation of step (1), each camera is calibrated in turn based on a standard dot calibration board to obtain the internal parameter matrix, external parameter matrix and distortion coefficient of each camera, so as to perform distortion correction on each camera. Specifically, for each camera, multiple calibration board pictures are collected from various angles, the corner coordinates of each calibration board are extracted using the findChessboardCorners function in opencv, and then the sub-pixel corner information is further extracted using cornerSubPix to reduce the calibration deviation. Then, the world coordinate system is set, the three-dimensional coordinates of the corners are calculated according to the calibration board parameters, and finally the calibrateCamera function is used for calibration to obtain the internal parameter matrix Ko, external parameter matrix Ro and distortion coefficient of each camera. Finally, the distortion correction function undistort in opencv is used to substitute the camera internal parameter matrix and distortion coefficient to achieve distortion correction of the collected images.
[0088] In the specific implementation of step (2), multiple images of each surface of the cake collected by multiple cameras after distortion correction are obtained, and several pairs of images to be stitched with overlapping regions are obtained. As Figure 2 shown, since the area of each surface of the cake is large, it is not possible to collect the whole with a single camera and the camera resolution requirements are relatively high. Therefore, two cameras are used to collect the upper and lower end faces of the cake 2, as Figure 2As shown, the multi-camera acquisition device for acquiring images of silk cakes mainly includes a silk cake support base 1, a circumferential surface acquisition camera 3, an upper end surface acquisition camera 4, and a lower end surface acquisition camera 5. There are two cameras for upper and lower end surface acquisition respectively, and 4 cameras for circumferential surface acquisition. Taking the left and right two silk cake images with overlapping areas and the 4 silk cake images on the circumferential surface obtained by the upper end surface acquisition camera 4 as an example, as Figure 3 shown.
[0089] In the specific embodiment of step (3), defect detection is performed on the silk cake images. If there are fluffs, flyings, broken wires, or looped wires on the surface of the silk cake, it is considered defective. If there are no defects, this detection is completed. Specifically, the YOLO algorithm trained with a large number of silk cake defect atlases is used for detection.
[0090] In the specific embodiment of step (4), the overlapping area of the image pair to be stitched is pre-estimated according to the camera parameters to obtain the proportion of the overlapping area in the image. Specifically, for the cameras that acquire the end surface images of the silk cake, their relative position relationship is translation, and the overlapping area satisfies the following calculation formula:
[0091]
[0092] where fov is the field of view angle of the camera, l is the distance between the two cameras, h is the distance from the camera to the surface of the silk cake, and φ is the proportion of the overlapping area of the silk cake end surface image in the original image area;
[0093] For the cameras that acquire the circumferential surface images of the silk cake, their relative position relationship is installed on a circle with a radius of r, where r is a value greater than the radius of the silk cake, and the overlapping area satisfies the following relationship:
[0094]
[0095] where λ is the proportion of the overlapping area of the silk cake circumferential surface image, fov is the size of the camera field of view angle, and β is the size of the angle between the two cameras.
[0096] In the specific embodiment of step (5), pixel adjustment and exposure compensation are performed on the image pair to be stitched, so that the images acquired by multiple cameras are more suitable for stitching and fusion. Specifically, preprocessing is performed on the problems of inconsistent scales of the same target in the images acquired by multiple cameras and color inconsistency caused by exposure differences between images;
[0097] Furthermore, the image pixel adjustment is to use pixel points with a value of 0 to perform edge filling on the images acquired by multiple cameras. Specifically, the upper, left, and lower three edges of the left image in each image pair to be stitched are filled, and the upper, right, and lower three edges of the right image are filled, so that the silk cake target is located in a relatively central position of the image, avoiding the phenomenon of target image loss during the subsequent perspective transformation process of the image pair to be stitched;
[0098] Furthermore, the exposure compensation is performed by using gain exposure. A gain coefficient is set for the images collected by multiple cameras to make the image intensities of the overlapping parts similar, which is mainly realized through the error function e. The empirical formula is as follows:
[0099]
[0100] where
[0101]
[0102] g i and g j are the gain coefficients of images i and j, R(i, j) is the overlapping part of the two images, R, G, B are the intensity values of each channel component of the image, N ij is the number of pixels in the overlapping part, σ N and σ g are the standard deviations of the error and the gain respectively. Usually, σ N = 10, σ g = 0.1, I ij is the average value of the image intensity in the overlapping part, n 1 and n 2 are the numbers of pixels in the width and height of the overlapping part respectively.
[0103] By taking the derivative of g with respect to e, the gain coefficient of each image is finally obtained.
[0104] In the specific embodiment of step (6), according to the proportion of the overlapping area in the image, the overlapping areas of the image pairs to be stitched processed in step (5) are extracted using corresponding proportion masks respectively. The improved LoFTR algorithm based on Transformer is used to extract feature points from the overlapping areas, and the feature point pairs with incorrect matches are removed according to the confidence of the feature point pairs, obtaining effectively matched feature point pairs. As shown in Figure 4 , the improved LoFTR algorithm is implemented through 4 modules: local feature CNN extraction module, feature transformation module, coarse-grained matching module, and fine-grained matching module.
[0105] Specifically, the implementation of the local feature CNN extraction module is as follows:
[0106] Using such as Figure 5The shown ResNetFPN network performs three downsamplings on the two input cake images to the levels of 1 / 2, 1 / 4, and 1 / 8 of the image size, and adds the scSE attention mechanism before each downsampling layer to facilitate the extraction of effective features. Then, bilinear interpolation is used to perform two upsamplings on the coarse-grained image feature map at the 1 / 8 level to obtain the feature maps at 1 / 4 and 1 / 2 levels in sequence, and fuse them with the high-level feature information. Finally, the coarse-grained image feature map at the 1 / 8 level is output and the fine-grained image feature map at 1 / 2 And adjust the activation function in the convolution process to LeakyReLU to avoid the dead point problem brought by Relu to a certain extent;
[0107] Specifically, as Figure 6 shown, the scSE attention mechanism consists of a spatial attention mechanism sSE and a channel attention cSE. Among them, the implementation process of sSE is to first use a 1×1×1 convolution on the input feature map to change its dimension from [C, H, W] to a feature map of [1, H, W], then use the sigmoid function for activation to obtain the spatial attention feature map, and then multiply the obtained result by the original feature map; among them, the implementation process of cSE is to transform the dimension of the input feature map from [C, H, W] to [C, 1, 1] through a global average pooling layer, and then use a C / 2×1×1 and a C×1×1 convolution in sequence for information processing to obtain a C-dimensional vector, then use the sigmoid function for normalization to obtain the corresponding mask, and finally multiply it by the original channel feature to obtain the channel-information calibrated feature map; scSE adds the results of cSE and sSE to obtain the final result, realizing the extraction of channel and spatial information.
[0108] Specifically, the implementation process of the feature conversion module is as follows:
[0109] For the incoming coarse-grained image feature map Use self-attention and cross-attention to extract local features related to position and context, and convert the features into a feature representation that is easy to match. The transformed feature representation is and
[0110] Specifically, the implementation steps of the coarse-grained matching module include:
[0111] Use a differentiable matching layer to match the features and to obtain a confidence matrix P c , and then use the mutual nearest neighbor algorithm to select matching pairs to obtain the coarse-grained matching prediction M c ;
[0112] Specifically, for each coarse-grained matching pair, the fineness matching module extracts a local window of size w×w from the fine-grained image feature map at the position, passes it into the self-attention and cross-attention layers, and uses softmax to calculate the score matrix and matching probability matrix of the window. Finally, the coordinates of the final feature matching points are obtained in the fine-grained image feature map and output.
[0113] The extracted feature matching point pairs are as Figure 7 shown, Figure 8 which is a feature point pair matching map after screening according to the confidence level.
[0114] In the specific implementation of step (7), based on the effective matching feature point pairs, the PROSAC algorithm is used for random sampling, and then the homography matrix of the image pairs to be stitched obtained in step (5) is solved. Specifically:
[0115] The progressive consistent sampling algorithm (PROSAC) is used for random sampling from the set of extracted feature points to reduce the computational amount and improve the detection speed.
[0116] The optimal homography transformation matrix between images with overlapping regions is defined as follows:
[0117]
[0118] And it satisfies:
[0119]
[0120] where (x i , y i ) are the coordinates of the feature points sampled in the left image in the overlapping region, and (x′ i , y′ i ) are the coordinates of the feature points sampled in the right image;
[0121] The homography transformation matrix is obtained to minimize the reprojection error rate, where the calculation method of the reprojection error rate is as follows:
[0122]
[0123] It should be noted that the four images to be stitched on the circumferential surface need to be stitched pairwise and passed in three times. Two images with overlapping regions are passed in each time for feature point extraction and transformation matrix solution.
[0124] In the specific implementation of step (8), the obtained homography matrix is used to perform perspective transformation on the right image to be stitched in the pair of images to be stitched processed in step (5), and the left image to be stitched is copied onto the perspective-transformed right image to be stitched, thus realizing direct stitching of the images. For direct stitching of images, perspective transformation needs to be carried out first. Specifically:
[0125] Use the homography matrix H obtained in step (7) to perform perspective transformation on the right image in the pair of images to be stitched, ensuring that the right image is transformed to the same plane as the left image, and its transformation process satisfies:
[0126]
[0127] Among them, (u, v) are the coordinates of the right image (the right image when not stitched together) for perspective transformation, and the coordinates of the image obtained after perspective transformation are (x, y). The calculation formula is as follows:
[0128]
[0129]
[0130] Then copy the left image in the image with overlapping regions onto the perspective-transformed right image to realize direct stitching of the images, as Figure 9 shown.
[0131] In the specific implementation of step (9), the weighted average image fusion algorithm is used to fuse the overlapping regions of the directly stitched images obtained in step (8) to eliminate the stitching seam and obtain an image of the complete cake surface without overlapping regions. Using the weighted average image fusion algorithm to fuse the overlapping regions of the directly stitched images in step (8) is mainly based on the processing of pixel points in this region. Specifically:
[0132]
[0133] Among them, g(x, y) is the pixel of the fused image point, g 1 (x, y) is the left image pixel point in the stitched image, g 2 (x, y) is the right image pixel point, g 1 ∩g 2 is the overlapping part in the stitched image, ω is the weight value, 0 < ω < 1, and pixel points are distributed by weighted average to achieve a seamless fusion effect. The final result is as Figure 10 shown.
[0134] In the specific implementation of step (10), the images of the non-repetitive regions on each surface are input into the host computer again for defect detection to determine the type, location, and number of defects, thereby rating the quality of the cake. The images of each surface of the fused cake are input into the host computer for defect detection. Specifically, the YOLO algorithm trained with a large number of cake defect atlases is used for detection, and the type, location, and number of defects are determined.
[0135] This application also compares the SURF algorithm and the improved LoFTR algorithm. The comparison is mainly carried out in terms of the number of feature points extracted, the total time consumed for image fusion, and the image fusion quality evaluation indicators PSNR and SSIM for the same spliced cake image. The results are shown in Table 1:
[0136] Table 1 Comparison table of SURF algorithm and improved LoFTR algorithm
[0137] Number of Feature Point Matches Total Time Consumed / s PSNR SSIM SURF Algorithm 1568 2.419 20.52 0.464 Improved LoFTR Algorithm 1972 0.476 20.65 0.558
[0138] From the above comparison, it can be seen that for the spliced cake image, the improved LoFTR algorithm has more advantages than the SURF algorithm in all aspects.
[0139] A comparison is also made between the LoFTR algorithm and the improved LoFTR algorithm. Mainly, 3 groups of cake images to be matched are selected for comparison in terms of the number of feature points extracted and the confidence range of feature point pairs. The results are shown in Table 2:
[0140] Table 2 Comparison table of LoFTR algorithm and improved LoFTR algorithm
[0141]
[0142] As can be seen from Table 2, for the three groups of cake images to be matched, the improved LoFTR algorithm is superior to the LoFTR algorithm in terms of the number of feature point pairs extracted and the confidence level, and is more suitable for application in image splicing and fusion work.
[0143] The present invention proposes a method for detecting cake defects based on the improved LoFTR algorithm for the problem of splicing cake images collected by multiple cameras, which solves the problems of insufficient extraction ability of the existing key point extraction algorithm for weakly textured images, repeated statistical defects due to incomplete collection by a single camera for large-sized cakes, overlapping image regions collected by multiple cameras, and long required time, enhances the key point extraction ability, and improves the splicing and fusion speed.
[0144] Corresponding to the embodiment of the method for detecting cake defects based on the improved LoFTR algorithm described above, this application also provides an embodiment of a seamless fusion device for cake images collected by multiple cameras based on the improved LoFTR algorithm.
[0145] Figure 11It is a block diagram of a seamless fusion device for multi-camera captured cake images based on an improved LoFTR algorithm shown according to an exemplary embodiment. Referring to Figure 11 , the device may include:
[0146] Calibration and correction module 101: Calibrate each camera in turn based on a standard dot calibration board to obtain the internal parameter matrix, external parameter matrix and distortion coefficient of each camera, so as to perform distortion correction on each camera;
[0147] Image acquisition module 102: Acquire multiple cake images on each surface of the cake captured by multiple cameras after distortion correction, and obtain several pairs of images to be stitched with overlapping regions;
[0148] First defect detection module 103: Detect defects in the cake images. If there are flyings, fuzz, broken wires, or looped wires on the cake surface, it means there are defects. If there are no defects, this detection is completed. If there are defects, continue to the next step;
[0149] Overlapping region ratio calculation module 104: Pre-estimate the overlapping regions of the pairs of images to be stitched according to the camera parameters, and obtain the ratio of the overlapping regions to the images;
[0150] Adjustment and compensation module 105: Perform pixel adjustment and exposure compensation on the pairs of images to be stitched to facilitate stitching and fusion;
[0151] Feature point extraction and pairing module 106: According to the ratio of the overlapping region to the image, use corresponding ratio masks to extract the overlapping regions of the pairs of images to be stitched processed by the adjustment and compensation module, use the improved LoFTR algorithm based on Transformer to extract feature points from the overlapping regions, and eliminate the wrongly matched feature point pairs according to the confidence of the feature point pairs to obtain effectively matched feature point pairs;
[0152] Matrix calculation module 107: Based on the effectively matched feature point pairs, use the PROSAC algorithm for random sampling, and then solve the homography matrix of the pairs of images to be stitched processed by the adjustment and compensation module;
[0153] Direct stitching module 108: Use the obtained homography matrix to perform perspective transformation on the right image to be stitched in the pairs of images to be stitched processed by the adjustment and compensation module, and copy the left image to be stitched onto the perspective-transformed right image to be stitched;
[0154] Image fusion module 109: Use a weighted average image fusion algorithm to fuse the overlapping regions of the directly stitched images obtained by the direct stitching module, eliminate the stitching seams, and obtain an image of the complete cake surface without overlapping regions;
[0155] Second defect detection module 110: Detect defects in the images of the complete and non-repetitive regions on each surface, determine the types, positions, and numbers of the defects, and thus rate the quality of the cake.
[0156] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0157] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0158] Correspondingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the cake defect detection method based on the improved LoFTR algorithm as described above. As Figure 12 shown, it is a hardware structure diagram of a device with any data processing ability where the cake defect detection method based on the improved LoFTR algorithm provided by the embodiment of the present invention is located. In addition to Figure 12 the processors, memory, and network interfaces shown, any device with data processing ability where the device in the embodiment is located usually also includes other hardware according to the actual functions of the device with any data processing ability, which will not be elaborated here.
[0159] Correspondingly, the present application further provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the method for detecting bobbin defects based on the improved LoFTR algorithm as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.
[0160] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the content disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0161] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for detecting wire cake defects based on an improved LoFTR algorithm. It is characterized in that include: (1) Calibrate each camera in turn based on the standard dot calibration plate to obtain the intrinsic parameter matrix, extrinsic parameter matrix and distortion coefficient of each camera, so as to perform distortion correction for each camera; (2) obtaining multiple images of the surfaces of the silk cakes collected by multiple cameras after distortion correction, and obtaining a plurality of image pairs to be stitched with repeated areas; (3) performing defect detection on the silk cake image. If there are hairy threads, hairiness, broken threads, or looped threads on the surface of the silk cake, the silk cake is defective. If there are no defects, the inspection is completed. If there are defects, proceed to the next step. (4) pre-estimating the overlapping area of the pair of images to be stitched according to camera parameters to obtain the proportion of the overlapping area to the image; (5) performing pixel adjustment and exposure compensation on the image pair to be stitched to facilitate stitching and fusion; (6) According to the ratio of the overlapping area to the image, use masks of corresponding ratios to extract the overlapping areas of the image pairs to be stitched after processing in step (5), use the improved Transformer-based LoFTR algorithm to extract feature points from the overlapping areas, and remove incorrectly matched feature point pairs according to the confidence of the feature point pairs to obtain effectively matched feature point pairs; (7) Based on the effectively matched feature point pairs, random sampling is performed using the PROSAC algorithm to solve the homography matrix of the image pair to be stitched obtained in step (5); (8) using the obtained homography matrix to perform perspective transformation on the right image to be stitched in the pair of images to be stitched obtained in step (5), and copying the left image to be stitched to the right image to be stitched after the perspective transformation; (9) using a weighted average image fusion algorithm to fuse the overlapping areas of the directly stitched images obtained in step (8) to eliminate the stitching seams and obtain an image of the silk cake surface that is complete and has no repeated areas; (10) performing defect detection on the images of the complete and non-repetitive areas of each surface to determine the type, location and number of defects, thereby rating the quality of the silk cake; In step (6), the improved Transformer-based LoFTR algorithm is implemented by a local feature CNN extraction module, a feature conversion module, a coarse-grained matching module, and a fine-grained matching module, wherein: The local feature CNN extraction module uses the ResNetFPN network to perform three downsamplings on the two input cake images to the levels of 1 / 2, 1 / 4, and 1 / 8 of the image size. The coarse-grained image feature map at the 1 / 8 level is upsampled twice using the bilinear interpolation method to obtain the feature maps at 1 / 4 and 1 / 2 levels in sequence, and is fused with the high-level feature information to output the coarse-grained image feature map at the 1 / 8 level and the fine-grained image feature map at 1 / 2 where an scSE attention mechanism is added before each downsampling; The feature transformation module is used to process the incoming coarse-grained image feature map Extract local features related to position and context using self-attention and cross-attention, and transform the features into a feature representation that is easy to match and The coarse-grained matching module uses a differentiable matching layer to match features and obtains a confidence matrix P c , uses the mutual nearest neighbor algorithm to select matching pairs, and obtains a coarse-grained matching prediction M c ; The fineness matching module for each coarse-grained matching pair extracts a local window of size w×w from the position in the fine-grained image feature map and feeds it into the self-attention and cross-attention layers, and uses softmax to calculate the score matrix and matching probability matrix of the window. Finally, the coordinates of the final feature matching points are obtained in the fine-grained image feature map and output 2. The method according to claim 1, It is characterized in that The multi-camera acquisition device for acquiring silk cake images includes a silk cake support seat, a circumferential surface acquisition camera, an upper end surface acquisition camera and a lower end surface acquisition camera. There are two upper end surface acquisition cameras and two lower end surface acquisition cameras, and four circumferential surface acquisition cameras.
3. The method according to claim 1, It is characterized in that In step (4), The repeated area between the images of the end faces of the silk cakes satisfies: Where fov is the field of view of the camera, l is the distance between the two cameras, h is the distance from the camera to the surface of the silk cake, and φ is the ratio of the overlapping area of the silk cake end surface image to the original image area; The repeated area between the images of the circumferential surface of the silk cake satisfies: where λ is the proportion of the overlapping area of the image on the circumferential surface of the bobbin, fov is the field of view angle of the camera, and β is the angle between the two cameras.
4. The method according to claim 1, characterized in that in step (5), the pixels are adjusted to use pixel points with a value of 0 to perform edge filling on the images collected by multiple cameras; the exposure compensation is performed by using gain exposure, and the gain coefficient is implemented through the error function e: where g i and g j are the gain coefficients of images i and j, R(i, j) is the overlapping part of the two images, R, G, B are the intensity values of each channel component of the image, N ij is the number of pixels in the overlapping part, σ N and σ g are the standard deviations of the error and the gain respectively. Usually, σ N = 10, σ g = 0.1, I ij is the average value of the image intensity in the overlapping part, n 1 and n 2 are the number of pixels in the width and height of the overlapping part respectively.
5. The method according to claim 1, characterized in that in step (7), the homography matrix H is calculated by the following formula and the matrix that minimizes the back-projection error rate is selected: Among them, (x i , y i ) are the coordinates of the feature points sampled from the left image to be stitched in the overlapping area, and (x' i , y' i ) are the coordinates of the feature points sampled from the right image to be stitched.
6. The method according to claim 1, characterized in that in step (8), for the coordinates (u, v) of the right image to be stitched that undergoes perspective transformation, the perspective transformation results in the image coordinates (x, y): [x′, y′, w′] = [u, v, w]H where H is the homography matrix.
7. The method according to claim 1, characterized in that in step (9), a weighted average image fusion algorithm is used to fuse the overlapping area of the directly stitched images in step (8), specifically: Among them, g(x, y) is the pixel of the fused image point, g 1 (x, y) is the pixel point of the left image to be stitched, g 2 (x, y) is the pixel point of the right image to be stitched, g 1 ∩g 2 is the overlapping part between the image pairs to be stitched, and ω is the weight value.
8. An electronic device, characterized in that comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
9. A computer-readable storage medium, on which computer instructions are stored, characterized in that when the instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Splicing method, system and equipment based on image overlapping region and medium
CN115526781A
Improved multi-scale eye fundus image splicing method based on Loftr algorithm
CN116152073A
Improved UDN joint-feature extraction-based pedestrian detection method
CN105335716A
Bill identification method and device
CN115187834A