Belt buckle surface anomaly detection algorithm based on machine vision technology

Through the belt buckle surface abnormality detection algorithm of machine vision and autoencoder network, the problems of low belt buckle detection efficiency and insufficient accuracy are solved, and efficient and accurate belt buckle abnormality detection is achieved, ensuring the safe and stable operation of the belt machine.

CN120298360AInactive Publication Date: 2025-07-11ANHUI TUOBANG CONVEYING EQUIP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510379288.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing belt buckle detection methods rely on manual inspection or simple sensors, which are inefficient and have limited accuracy, making it difficult to meet the requirements of modern industrial production for efficiency and reliability, and face the problems of scarce abnormal samples, complex visual environment and blurred target areas.

Method used

The belt buckle surface abnormality detection algorithm based on machine vision is used to capture the belt buckle area through the differential method, and image reconstruction and abnormality determination are carried out in combination with the autoencoder network, and accurate judgment is made using OTSU method and image morphology technology.

Benefits of technology

Real-time online detection of belt buckles is realized, the detection efficiency and accuracy are improved, the false alarm rate is reduced, and the safe and stable operation of the belt machine is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298360A_ABST
    Figure CN120298360A_ABST
Patent Text Reader

Abstract

The invention discloses a belt buckle surface anomaly detection algorithm based on a machine vision technology, and relates to the technical field of computer vision and image processing, the belt buckle area is accurately reconstructed, and the image before and after reconstruction is meticulously analyzed, so that a tiny abnormal area can be accurately identified; in the abnormal contrastive analysis, the OTSU method binarization processing and the image morphology technology are adopted, contour detection and area calculation are combined, the size of an abnormal area can be accurately quantified, when the area of a suspicious abnormal point exceeds 10% of the area of the belt buckle, it can be judged that the belt buckle is abnormal in time and accurately, the false alarm rate is greatly reduced, and the safety of the belt buckle is improved. And a powerful guarantee is provided for safe and stable operation of the belt conveyor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and image processing, and specifically to a surface anomaly detection algorithm for belt buckles based on machine vision technology. Background Art

[0002] The belt buckle is a key component in a belt conveyor, and its main function is to connect the belt into a closed loop. In actual use, the belt buckle is prone to damage such as broken teeth due to uneven stress, wear, or quality defects; such damage may cause the overall failure of the belt conveyor and result in serious consequences.

[0003] Existing detection methods mainly rely on manual inspection or simple sensors. This method is not only inefficient but also limited in accuracy, and it is difficult to meet the requirements of high efficiency and reliability in actual industrial production; therefore, developing an automated detection method based on machine vision to achieve efficient and accurate belt buckle defect detection has important research and application value.

[0004] In recent years, various anomaly detection methods for key components based on machine vision have been proposed. However, the detection of belt buckles still faces the following main challenges:

[0005] (1) Scarce anomaly samples: Although belt buckles are prone to failures during use, abnormal problems such as broken teeth are relatively rare, and the types of abnormal problems are diverse, making it difficult to form a large-scale dataset, thus restricting the development and optimization of machine vision algorithms.

[0006] (2) Complex visual environment: Belt conveyors usually operate in a dim and humid environment, and the surface of the belt buckle is easily disturbed by shadows or stains; these factors make the vision-based detection task more complex.

[0007] (3) Blurred target area: During the high-speed operation of the belt buckle, the captured images may be blurred or distorted. In addition, the target area of the belt buckle is not prominent, bringing additional difficulties to anomaly detection. Summary of the Invention

[0008] Aiming at the deficiencies of the existing technology, the present invention provides a surface anomaly detection algorithm for belt buckles based on machine vision technology, which can capture the belt buckle area during the operation of the belt conveyor and detect and analyze the surface anomalies; the system compares the size of the abnormal area with a preset threshold and feeds the detection result back to the control center to assist in decision-making.

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions: A surface anomaly detection algorithm for belt buckles based on machine vision technology, including the following steps:

[0010] Step 1. According to the machine vision device, visually monitor the belt during operation. Use the difference method to subtract the images of adjacent frames to determine the representative feature values of the current frame image. From several groups of representative feature values, select the image to be detected. The specific sub-steps are as follows:

[0011] S11. From several consecutive images monitored by the machine vision device, identify the feature differences between adjacent frame images: confirm the image area associated with the belt within a preset pixel value range, and the pixel value range is the preset range. Combine the edge contour of the image area with a two-dimensional coordinate system, where the horizontal axis of the two-dimensional coordinate system is combined with the horizontal contour line of the image area, and the vertical axis is combined with the vertical contour line of the image area;

[0012] S12. Use the coordinate points within the horizontal axis as feature points, confirm the feature perpendicular lines perpendicular to these feature points and located within the image area, sum the pixel values of several pixel points associated with the feature perpendicular lines to confirm the total feature of the perpendicular lines associated with this feature perpendicular line, and process them in sequence. Sort the different total features of the perpendicular lines corresponding to different feature points to confirm the sorting sequence of the total features of the perpendicular lines;

[0013] S13. Subtract the sorting sequences of the total features associated with adjacent frame images: subtract the total features of the perpendicular lines belonging to the same feature point to confirm the same-position difference values, and then sum several groups of the same-position difference values to lock the representative feature value of the latter frame image;

[0014] S14. Define a set of monitoring periods, and the monitoring period is the preset period. Determine the representative feature values associated with several frames of images within the preset period in sequence, and select the image associated with the maximum representative feature value as the image to be detected. If there are multiple groups of images to be detected, randomly select one group as the selected image to be detected;

[0015] Step 2. Use the binary processing method to process the image to be detected and lock the belt buckle area existing within the image to be detected: based on the set threshold, label the pixel points with pixel values lower than this threshold as 0-value points, and label the pixel points not lower than this threshold as 1-value points. The threshold is the preset value, and label the area covered by the 1-value points as the belt buckle area;

[0016] Step 3. Visually correct the determined belt buckle area. By calibrating the positions of its four vertices in the image and using a visual correction algorithm, calculate the correction matrix and use this matrix to complete the correction of the picture of the belt buckle area. The specific method is as follows:

[0017] S31. Place a 1m x 1m square grid in the image and determine the pixel coordinates of the four vertices of this square grid in the image through an image processing algorithm;

[0018] S32. Given the coordinates of the square grid in the real world and the pixel coordinates in the image, use the perspective transformation algorithm to calculate the perspective transformation matrix from the image coordinates to the real world coordinates;

[0019] S33. Use the calculated correction matrix to perform perspective transformation on the image containing the buckle area;

[0020] Step Four: Evenly divide the buckle area after visual correction. Select a feature edge from the contour line of the buckle area, and based on this feature edge, divide this buckle area into N rectangular areas. The specific sub-steps are as follows:

[0021] Identify the edge parallel to the Y-axis of the buckle area and mark it as the feature edge. Evenly divide this feature edge so that the feature edge is divided into N edge segments, where N is a preset value. Then, connect the edge segments in the same sorting position from top to bottom to evenly divide the buckle area into N rectangular areas, and perform pixel scaling on each rectangular area so that each rectangular area is scaled to an adjusted area of 64x64;

[0022] Step Five: From the buckle areas associated with 20 preset images, confirm 20×N images, and use this as a training set to train the autoencoder network. The specific method is as follows:

[0023] Construct an autoencoder network that includes 3 convolutional layers, a 128-dimensional encoding layer, and 3 pooling layers; the encoding layer compresses the input 64x64 pixel image into a 128-dimensional low-dimensional representation, and the decoding layer reconstructs the low-dimensional representation into a 64x64 pixel image;

[0024] Use the training set to train the autoencoder. The goal is to minimize the error between the reconstructed image and the original input image; use the mean squared error (MSE) as the loss function, and then use the stochastic gradient descent (SGD) or Adam optimizer to update the parameters;

[0025] Step Six: Pass the several 64x64 adjusted areas associated with the buckle area through the autoencoder network for low-dimensional representation, and then perform image reconstruction through the decoding layer. Stitch the reconstructed adjusted areas together in the original order to form a complete reconstructed image. The specific method is as follows:

[0026] Send the adjusted areas to the encoding layer of the autoencoder network so that each adjusted area is compressed into a 128-dimensional low-dimensional representation;

[0027] Input the 128-dimensional low-dimensional representation into the decoding layer, and through a series of deconvolution and upsampling operations, reconstruct it into a 64x64 image:

[0028] Map the 128-dimensional low-dimensional representation to a high-dimensional vector through a fully connected layer;

[0029] The image is blurred by the upsampling method, and a learnable convolution kernel slides on the image to perform weighted summation on local regions, extract and integrate feature information, and fill in the details lost after upsampling;

[0030] After multiple upsampling and convolution operations, the size of the image gradually returns to 64x64. The output layer then uses an activation function (such as Sigmoid) to map the pixel values to the range [0,1], making them conform to the representation range of image pixel values, and finally obtaining a reconstructed 64x64 pixel image to complete the reconstruction process;

[0031] Step Seven: Confirm the pixel difference between the buckle area and the reconstructed image, lock the abnormal difference area and the detection result, and perform relevant display. The specific method is as follows:

[0032] Calculate the absolute value of the difference between the corresponding pixel values of the original image and the reconstructed image to obtain a difference image;

[0033] Use the OTSU method to perform binary processing on the difference image, automatically determine a suitable threshold, and convert the difference image into a binary image. Among them, white pixels represent areas where abnormalities may exist, and black pixels represent normal areas. Pixels with gray values higher than the threshold are set to white, and pixels lower than the threshold are set to black;

[0034] Apply the erosion and dilation techniques in image morphology to process the binary image, use a 5x5 all-ones matrix as the structuring element to remove noise points and fill holes;

[0035] Use a contour detection algorithm to find the contours of all white areas in the binary image and calculate the area enclosed by each contour;

[0036] Find the white area with the largest area and calculate the proportion of its area to the total area of the buckle; when this proportion exceeds 10%, it is determined that there is an abnormality in the buckle and an alarm message is sent.

[0037] The present invention provides a buckle surface abnormality detection algorithm based on machine vision technology. Compared with the prior art, it has the following beneficial effects:

[0038] The video stream image capture technology based on the difference method of the present invention can quickly process a large amount of picture information, automatically capture images containing buckles without manual intervention. At the same time, the entire detection process is highly automated, and from image preprocessing, feature extraction to abnormality determination, it can be quickly completed. Compared with manual detection, the detection efficiency has been improved by an order of magnitude, and real-time online detection of buckles can be achieved, meeting the urgent need for high efficiency in modern industrial production, effectively reducing production downtime, and improving overall production efficiency;

[0039] The neural network adopted shows strong robustness in dealing with such interference factors. Through the buckle area recognition technology based on semantic segmentation, the buckle area can be accurately recognized under complex backgrounds, reducing the influence of interference factors on the detection results; the abnormal determination method based on autoencoders and digital image processing can also effectively eliminate environmental interference when reconstructing images and analyzing abnormalities, accurately judging whether there are abnormalities in the buckle, improving the reliability and accuracy of the detection results;

[0040] By precisely reconstructing the buckle area and carefully analyzing the images before and after reconstruction, tiny abnormal areas can be accurately identified; in the abnormal comparative analysis, the OTSU method for binary processing and image morphology technology are adopted, combined with contour detection and area calculation, to accurately quantify the size of the abnormal area. When the area of the suspicious abnormal point exceeds 10% of the buckle area, it can timely and accurately judge that there are abnormalities in the buckle, greatly reducing the false alarm rate and providing a strong guarantee for the safe and stable operation of the belt conveyor. Brief Description of the Drawings

[0041] Figure 1 It is a schematic flowchart of the method of the present invention. Detailed Embodiments

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] The First Embodiment

[0044] Please refer to Figure 1 , this application provides a buckle surface abnormal detection algorithm based on machine vision technology, including the following steps:

[0045] Step 1: According to the machine vision device, visually monitor the running belt, use the difference method to subtract the images of adjacent frames before and after, determine the representative feature values of the current frame image, and select the image to be detected from several groups of representative feature values. Specifically, the belt is in a rotating state, and the vision device is located at the center position above the belt to obtain and analyze the images of the outer surface of the belt. There are relevant differences in the pixel values of each frame of the image. For the specific image where the buckle exists, the degree of difference is the largest. Therefore, this method can be used to determine the image to be detected with the most obvious buckle features. The specific sub-steps for selecting the image to be detected are as follows:

[0046] S11. Identify the feature differences between adjacent frame images from several consecutive images monitored by the machine vision device: Confirm the image area associated with the belt within a preset pixel value range. The pixel value range is a preset range determined in advance by the operator and is used to extract the area including the belt (including the belt buckle). Combine the edge contour of the image area with a two-dimensional coordinate system. The horizontal axis of the two-dimensional coordinate system is combined with the horizontal contour line of the image area, and the vertical axis is combined with the vertical contour line of the image area (since there are four sets of contour lines, only two sets of edge contour lines are selected here).

[0047] S12. Use the coordinate points on the horizontal axis as feature points, confirm the feature perpendicular lines perpendicular to these feature points and located within the image area, sum the pixel values of several pixel points associated with the feature perpendicular lines, confirm the total feature of the perpendicular lines associated with these feature perpendicular lines, and process them in sequence. Sort the different total feature values of the perpendicular lines corresponding to different feature points to confirm the sorting sequence of the total feature values of the perpendicular lines.

[0048] S13. Perform a subtraction operation on the sorting sequences of the total feature values associated with adjacent frame images: Subtract the total feature values of the perpendicular lines belonging to the same feature point to confirm the same-position difference values, and then sum several groups of the same-position difference values to lock the representative feature value of the latter frame image. For example, assume that the sorting sequence of the total feature values associated with the previous frame image of adjacent frame images is {T1, T2, T3, ……, Tn}, and the sorting sequence of the total feature values associated with the latter frame image is {Q1, Q2, Q3, ……, Qn}. Use: Representative feature value = |T1 - Q1| + |T2 - Q2| + |T3 - Q3| + …… + |Tn - Qn| to confirm the representative feature value associated with the latter frame image.

[0049] S14. Define a set of monitoring periods. The monitoring period is a preset period (the preset period is determined according to the rotation period of the corresponding pulley). Determine the representative feature values associated with several frame images within the preset period in sequence, and select the image associated with the maximum representative feature value as the image to be detected. If there are multiple groups of images to be detected (that is, there are multiple groups of images with the same representative feature value), randomly select one group as the selected image to be detected.

[0050] Specifically, by using the method of subtracting the previous image from the subsequent image, analyze whether the belt buckle is passing through the camera. When the belt buckle passes through, the difference between the previous image and the subsequent image is relatively large. We sum along the Y-axis of the image and select the maximum value as the representative feature value of the current image, which we call Peak. When Peak exceeds a certain threshold, we start recording images and select the image with the maximum Peak in one of the recorded image segments as the target image.

[0051] Step 2: Process the image to be detected using binaryzation to lock the buckle area existing in the image to be detected. Based on the set threshold, the pixel points with pixel values lower than this threshold are marked as 0-value points, and the pixel points not lower than this threshold are marked as 1-value points. The threshold is a preset value determined by the operator according to experience, generally between the pixel values of the belt surface points and the buckle points. The area covered by the 1-value points is marked as the buckle area.

[0052] Step 3: Perform visual correction on the determined buckle area. By calibrating the positions of its four vertices in the image and using a visual correction algorithm, calculate the correction matrix and use this matrix to complete the correction of the buckle area image. The specific correction method is as follows:

[0053] S31: Place a 1m x 1m square grid in the image and determine the pixel coordinates of the four vertices of the square grid in the image through image processing algorithms (such as edge detection and corner detection). Common corner detection algorithms include Harris corner detection, Shi-Tomasi corner detection, etc. Through these algorithms, the specific positions of the four vertices of the square grid in the image can be accurately found, providing basic data for subsequent calculation of the correction matrix.

[0054] Harris corner detection determines corners based on the change in pixel gray values within a local window of the image. When a small window slides on the image, if the pixel gray values within the window change significantly in both the x and y directions, then this position may be a corner. If there is only a significant change in one direction, it may be an edge. If the changes in both directions are small, it is a flat area.

[0055] Confirm the gradient Tx in the X direction and the gradient Ty in the Y direction associated with each pixel point in the image. Based on the gradients confirmed at the corresponding positions, confirm the structure tensor M, which is a 2×2 matrix. where ∑ w represents the summation in a small window centered on the current pixel. Calculate the corner response value R based on the structure tensor.

[0056] Its R = det(M) - k(trace(M)) 2 , where det(M) is the determinant of matrix M, trace(M) is the trace of matrix M, and k is an empirical constant, usually taking values between 0.04 - 0.06.

[0057] Then set a threshold. When the response value R of a pixel point is greater than this threshold, this pixel point is recognized as a corner. At the same time, to avoid repeated detection of corners, the non-maximum suppression method is usually used, that is, only the corner with the largest response value is retained, and the points with smaller response values around are suppressed.

[0058] S32. Given the coordinates of a square grid in the real world (assuming the four vertex coordinates are (0, 0), (0, 1), (1, 1), (1, 0)) and the pixel coordinates in the image, use the perspective transformation algorithm to calculate the perspective transformation matrix from the image coordinates to the real world coordinates. Taking the cv2.getPerspectiveTransform function in OpenCV as an example, it receives the source coordinates and target coordinates (real world coordinates) of the four vertices of the square grid in the image. Through the internal calculation logic of the algorithm, it outputs a 3x3 perspective transformation matrix. This matrix contains the transformation information of the image in aspects such as translation, rotation, scaling, and distortion, and is used to accurately correct the image subsequently. The cv2.getPerspectiveTransform function performs a series of complex mathematical operations on the input source coordinates and target coordinates, and finally outputs a 3x3 perspective transformation matrix. Its core steps include establishing a system of equations, solving the system of linear equations, and constructing the transformation matrix:

[0059] Establish a system of equations: In perspective transformation, a point (x, y) in the source image is mapped to a point (x′, y′) in the target image after transformation. The transformation relationship can be represented in homogeneous coordinates as:

[0060] For the four vertices of the square grid, such a transformation relationship can be established for each vertex. Since in homogeneous coordinates two equations can be obtained for each vertex, and 8 equations can be obtained for the four vertices;

[0061] The 8 equations form a system of linear equations containing 9 unknowns (a11, a12, ……, a33). Since the number of unknowns is 1 more than the number of equations, it is necessary to utilize the characteristics of homogeneous coordinates. Usually, let a33 = 1, so that the problem is transformed into solving a system of linear equations with 8 unknowns. The cv2.getPerspectiveTransform function internally uses specific linear algebra algorithms (such as Gaussian elimination method, etc.) to solve this system of equations and obtain the values of a11, a12, ……, a32;

[0062] Substitute the values of the 8 unknowns obtained by solving into the 3x3 matrix This matrix is the finally output perspective transformation matrix, which can perform perspective transformation on the square grid in the source image (determined by the vertex coordinates) according to the requirements of the target coordinates (real world coordinates), and realize the conversion from image coordinates to real world coordinates;

[0063] S33. Use the calculated correction matrix to perform a perspective transformation on the image containing the buckle area; in OpenCV, this operation can be achieved using the cv2.warpPerspective function. This function takes the input image, the correction matrix, and the size of the output image as parameters. By calculating the transformation of the coordinates of each pixel in the image, the buckle area in the image is transformed from the tilted and deformed state during shooting to a frontal view angle, making the shape and size of the buckle area closer to the real situation, effectively reducing the deformation caused by the shooting angle, and providing more accurate image data for subsequent operations such as buckle segmentation, feature extraction, and anomaly detection. The specific way of the change is as follows:

[0064] Use the cv2.imread function to read the image, ensuring that the image format is correct and the path is error-free;

[0065] Specify the size of the output image as needed, usually the same as the size of the original image or adjusted according to specific requirements;

[0066] Pass in the original image, the correction matrix, and the size of the output image; perform a perspective transformation on the original image according to the correction matrix to generate the corrected image;

[0067] Use the cv2.imshow function to display the corrected image, or use the cv2.imwrite function to save it to the specified path.

[0068] Step Four. Perform an equal division process on the buckle area that has completed visual correction. Select a feature edge from the contour line of the buckle area, and based on this feature edge, divide this buckle area into N rectangular areas. The specific sub-steps of the division are as follows:

[0069] S41. Confirm the edge parallel to the Y-axis of the buckle area and record it as the feature edge. Perform an equal division process on this feature edge so that the feature edge is divided into N edge segments, where N is a preset value. And according to the top-down method, connect the edge segments at the same sorting position to divide the buckle area into N rectangular areas with the same area. Specifically, the corresponding buckle area is in a strip shape, and both long sides of the strip shape are equally divided. During the equal division process, connect the corresponding equal division points. From top to bottom, connect the equal division points at the same position, and the corresponding buckle area can be effectively divided into N rectangular areas. Then perform pixel scaling on each rectangular area so that each rectangular area is scaled to an adjusted area of 64x64;

[0070] Step Five. Confirm 20×N pictures from the buckle areas associated with 20 preset images, and use this as the training set to train the autoencoder network:

[0071] Construct an autoencoder network that includes 3 convolutional layers, a 128 - dimensional encoding layer, and 3 pooling layers; the encoding layer compresses the input 64x64 pixel image into a 128 - dimensional low - dimensional representation, and the decoding layer reconstructs the low - dimensional representation into a 64x64 pixel image:

[0072] The fully - connected layer receives the 128 - dimensional low - dimensional vector and maps it to a higher dimension (such as 8192 dimensions) through a weight matrix; the output dimension is set according to the final reconstructed size and number of channels of the image. The fully - connected layer integrates and transforms the feature information in the low - dimensional vector, distributes the information over a wider dimension, and prepares for subsequent reshaping and transposed convolution operations; it converts the abstract features in the low - dimensional representation into a form more suitable for image reconstruction, increasing the diversity and availability of features;

[0073] Reshape the data output by the fully - connected layer into a specific shape (such as (8, 8, 128)) through a Reshape operation; the determination of this shape is closely related to the subsequent transposed convolution operation, and it is the starting structure for upsampling and convolution operations; the reshape operation converts one - dimensional data into three - dimensional data, constructs the spatial structure prototype of the image, and lays the foundation for restoring the image size and details;

[0074] The upsampling layer (such as UpSampling2D((2, 2))) magnifies the image size in the height and width directions by replicating pixels; after each upsampling, the image size doubles, gradually restoring from a smaller size to the target 64x64 pixels; however, simple upsampling will blur the image and lose details;

[0075] The convolutional layer performs a convolution operation on the image after each upsampling; the convolution kernel slides on the image, performs weighted summation on the surrounding pixels, and extracts and learns image features; the convolution operation can capture local patterns in the image, such as edge, texture and other information, and integrate these features into the magnified image, filling the detail loss caused by upsampling, making the image closer to the original image;

[0076] The output of the last convolutional layer is the reconstructed image. The number of convolution kernels in the output layer corresponds to the number of channels of the image (3 channels for color images), and the feature information processed previously is converted into pixel values of the image through convolution operations; the sigmoid activation function is used to map the output values to the 0 - 1 interval, which conforms to the representation range of image pixel values, thereby generating the final 64x64 pixel reconstructed image;

[0077] Use the training set to train the autoencoder, with the goal of minimizing the error between the reconstructed image and the original input image; the mean squared error (MSE) can be used as the loss function, and the Stochastic Gradient Descent (SGD) or Adam optimizer can be used for parameter update:

[0078] The mean squared error is used to quantify the difference between the original input image and the reconstructed image. For a batch of training images, assume the i-th original image is X i , and the corresponding reconstructed image is The calculation formula for the MSE loss L of this batch of images is: where n is the data of this batch of training images, and ||·|| 2 represents the squared Euclidean norm, that is, the sum of the squares of the corresponding pixel differences; this loss value reflects the similarity between the current reconstructed image of the autoencoder and the original image. The smaller the value, the better the reconstruction effect. For example, if the value of a pixel in the original image is 0.8 and the value in the reconstructed image is 0.6, the contribution of this pixel to the loss is (0.8 - 0.6) 2 = 0.04; by accumulating and averaging the contributions of all pixel points, the loss of the entire image is obtained;

[0079] SGD updates the parameters based on the gradient of the loss function with respect to the autoencoder parameters (such as the weights and biases of the convolutional layer, the weights and biases of the fully connected layer, etc.); for the above MSE loss function, the chain rule is used to calculate its gradient with respect to each parameter. Assume the parameter is w and the loss function is L, then the gradient represents the rate of change of the loss function L with respect to the parameter w. For example, for a weight w ij in the convolutional layer, through the chain rule, propagating from the loss direction of the output layer, the

[0080] is calculated. Then the parameter update is performed: after calculating the gradient, the parameters are updated according to the following formula:

[0081] where η is the learning rate, which is a hyperparameter that determines the step size of each parameter update. If the learning rate is set too large, the parameter update may skip the optimal solution, resulting in the loss function not decreasing but increasing. If the learning rate is set too small, the parameter update speed will be very slow and the training time will be greatly extended. For example, if the current weight w = 0.5 and the calculated gradient and the learning rate η = 0.01, then the updated weight w = 0.5 - 0.01×0.2 = 0.498; SGD uses a small batch of training samples (instead of the entire training set) each time to calculate the gradient and update the parameters, which makes the calculation more efficient and helps to jump out of the local optimal solution;

[0082] Step 6. Perform image reconstruction: Use the trained autoencoder network to reconstruct the buckle area. Represent the several 64x64 adjustment areas associated with the buckle area in a low-dimensional manner through the autoencoder network, and then perform image reconstruction through the decoding layer. Stitch the reconstructed adjustment areas together in the original order to form a complete reconstructed image;

[0083] The specific method for image reconstruction is as follows: The adjusted region is sent to the encoding layer of the autoencoder network, and each adjusted region is compressed into a 128-dimensional low-dimensional representation, which contains the key feature information of the original image;

[0084] The 128-dimensional low-dimensional representation is input into the decoding layer, and after a series of deconvolution and upsampling operations, it is reconstructed into a 64x64 image:

[0085] The 128-dimensional low-dimensional representation is mapped to a higher-dimensional vector through a fully connected layer. Suppose the 128-dimensional vector is mapped to an 8192-dimensional vector. Here, 8192 is determined according to the size and number of channels of the image to be reconstructed later, and it can be reshaped into a three-dimensional tensor for subsequent convolution operations;

[0086] The image is blurred by the upsampling method. The learnable convolution kernel slides on the image, performs weighted summation on the local region, extracts and integrates feature information, fills in the details lost after upsampling, and makes the image clearer and more realistic. At the same time, the activation function (such as ReLU) in the convolutional layer introduces non-linearity, increasing the expressive ability of the model;

[0087] After multiple upsampling and convolution operations, the size of the image gradually returns to 64x64, but the range of the output pixel values may not be between [0,1]. The output layer uses an activation function (such as Sigmoid) to map the pixel values to the interval [0,1], making it conform to the representation range of image pixel values, and finally obtaining a reconstructed 64x64 pixel image to complete the reconstruction process.

[0088] Step Seven: Confirm the pixel difference between the buckle region and the reconstructed image, lock the abnormal difference region and the detection result, and perform relevant display. The specific method for locking is as follows:

[0089] Calculate the absolute value of the difference between the corresponding pixel values of the original image and the reconstructed image to obtain a difference image. The difference image can be understood as a partial region image with pixel differences. By calculating the absolute value of the difference between their corresponding pixel values, the part of the original image that is different from the normal state can be highlighted. If there are abnormalities in the buckle, the pixel values of these abnormal regions in the original image and the reconstructed image will have obvious differences, and will be manifested as larger pixel values in the difference image; while the pixel value differences in the normal regions are smaller and close to 0 in the difference image;

[0090] Use the OTSU method to perform binary processing on the difference image, automatically determine a suitable threshold, and convert the difference image into a binary image, where white pixels represent the regions that may have abnormalities, and black pixels represent normal regions. The pixels with gray values higher than the threshold are set to white (indicating possible abnormalities), and the pixels lower than the threshold are set to black (indicating normal);

[0091] The erosion and dilation techniques in image morphology are used to process binary images. A 5x5 all-ones matrix is used as the structuring element. The erosion operation can remove small noise points. The dilation operation can fill small holes, making the abnormal area more continuous and complete. If most of the pixels around a certain pixel in the white area are black (i.e., do not conform to the shape of the structuring element), then this pixel will be set to black, which can remove small noise points because noise points are usually isolated and small-sized white pixels. The dilation operation is the opposite. It expands the white area, converts the black pixels around the white area into white according to the shape of the structuring element, fills small holes, and makes the abnormal area more continuous and complete, facilitating subsequent contour detection and analysis;

[0092] Use a contour detection algorithm (such as the cv2.findContours function in OpenCV) to find the contours of all white areas in the binary image and calculate the area enclosed by each contour. By finding the contours, each possible abnormal area can be treated as an independent object; calculating the area enclosed by each contour can quantify the size of the abnormal area and provide a numerical basis for subsequent abnormal judgment;

[0093] Find the white area with the largest area and calculate the ratio of its area to the total area of the belt buckle; when this ratio exceeds 10%, it is determined that the belt buckle is abnormal and an alarm message is issued.

[0094] Some of the data in the above formula are numerically calculated after removing their dimensions, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0095] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. Belt buckle surface anomaly detection algorithm based on machine vision technology, characterized in that It includes the following steps: Based on the visual monitoring of the belt during operation by a machine vision device, the difference method is used to calculate the difference between adjacent frames of images. The image to be detected is selected, and then the binaryzation processing method is used to lock the buckle area in the image to be detected. Visual correction is performed on the buckle area, and then the buckle area after visual correction is evenly divided into N rectangular areas; The training set is confirmed and the autoencoder network is trained. A number of 64x64 adjusted areas associated with the buckle area are represented in a low dimension through the autoencoder network, and then image reconstruction is performed through the decoding layer. The reconstructed adjusted areas are spliced into a complete reconstructed image in the original order; The pixel difference between the buckle area and the reconstructed image is confirmed to lock the abnormal difference area and the detection result.

2. The belt buckle surface anomaly detection algorithm based on machine vision technology according to claim 1, characterized in that, The specific sub-steps for selecting the image to be detected are as follows: S11. From several consecutive images monitored by the machine vision device, identify the feature differences between adjacent frame images: confirm the image area associated with the belt within a preset pixel value range, and the pixel value range is the preset range. Combine the edge contour of the image area with a two-dimensional coordinate system, where the horizontal axis of the two-dimensional coordinate system is combined with the horizontal contour line of the image area, and the vertical axis is combined with the vertical contour line of the image area; S12. Using the coordinate points within the horizontal axis as feature points, confirm the feature perpendicular lines perpendicular to these feature points and located within the image area. Sum the pixel values of several pixel points associated with the feature perpendicular lines to confirm the total feature of the perpendicular lines associated with these feature perpendicular lines, and process them in sequence. Sort the different total feature values of the perpendicular lines corresponding to different feature points to confirm the sorting sequence of the total feature values of the perpendicular lines; S13. Perform a difference operation on the sorting sequences of the total feature values associated with adjacent frame images: calculate the difference between the total feature values of the perpendicular lines belonging to the same feature point to confirm the same-position difference value, and then sum several groups of same-position difference values to lock the representative feature value of the latter frame image; S14. Define a set of monitoring periods, and the monitoring period is the preset period. Determine the representative feature values associated with several images within the preset period in sequence, and select the image associated with the maximum representative feature value as the image to be detected. If there are multiple groups of images to be detected, randomly select one group as the selected image to be detected.

3. The belt buckle surface anomaly detection algorithm based on machine vision technology according to claim 1, characterized in that, The binaryzation processing method includes: based on the set threshold, label the pixel points with pixel values lower than this threshold as 0-value points, and label the pixel points not lower than this threshold as 1-value points. The threshold is the preset value, and the area covered by the 1-value points is labeled as the buckle area.

4. The belt buckle surface anomaly detection algorithm based on machine vision technology according to claim 1, characterized in that, The specific method for correcting the buckle area is as follows: S31. Place a 1mx1m square grid in the image, and determine the pixel coordinates of the four vertices of the square grid in the image through an image processing algorithm; S32. Given the coordinates of the square grid in the real world and its pixel coordinates in the image, use the perspective transformation algorithm to calculate the perspective transformation matrix from the image coordinates to the real world coordinates; S33. Use the calculated correction matrix to perform perspective transformation on the image containing the buckle area.

5. The surface anomaly detection algorithm for belt buckles based on machine vision technology according to claim 1, wherein, The specific sub - steps for dividing the buckle area are as follows: Identify the side of the buckle area parallel to the Y - axis of the two - dimensional coordinate system and denote it as the characteristic side. Divide this characteristic side evenly so that it is divided into N side segments, where N is a preset value. Then, connect the side segments in the same sorting position from top to bottom, dividing the buckle area into N rectangular areas. Next, perform pixel scaling on each rectangular area so that each rectangular area is scaled to an adjusted area of 64x64.

6. The surface abnormality detection algorithm for belt buckles based on machine vision technology according to claim 1, characterized in that The training set includes: Identify 20×N pictures from the buckle areas associated with 20 preset images.

7. The surface abnormality detection algorithm for belt buckles based on machine vision technology according to claim 6, characterized in that, The specific method for training the auto - encoder network based on the training set is as follows: Construct an auto - encoder network that includes 3 convolutional layers, a 128 - dimensional encoding layer, and 3 pooling layers. The encoding layer compresses the input 64x64 - pixel image into a 128 - dimensional low - dimensional representation, and the decoding layer reconstructs the low - dimensional representation into a 64x64 - pixel image. Use the training set to train the auto - encoder. The goal is to minimize the error between the reconstructed image and the original input image. Use the mean squared error (MSE) as the loss function, and then use the stochastic gradient descent (SGD) or Adam optimizer to update the parameters.

8. The belt buckle surface anomaly detection algorithm based on machine vision technology according to claim 1, characterized in that, The specific method for image reconstruction through the decoding layer is as follows: Send the adjusted area to the encoding layer of the auto - encoder network so that each adjusted area is compressed into a 128 - dimensional low - dimensional representation. Input the 128 - dimensional low - dimensional representation into the decoding layer. After a series of de - convolution and up - sampling operations, it is reconstructed into a 64x64 image: Map the 128 - dimensional low - dimensional representation to a high - dimensional vector through a fully - connected layer; Blur the image through the up - sampling method. Slide a learnable convolutional kernel over the image, perform weighted summation on local areas, extract and integrate feature information, and fill in the details lost after up - sampling. After multiple up - sampling and convolution operations, the size of the image gradually returns to 64x64. The output layer then uses an activation function to map the pixel values to the range [0,1] to conform to the representation range of image pixel values, and finally obtains the reconstructed 64x64 - pixel image, completing the reconstruction process.

9. The surface anomaly detection algorithm for belt buckles based on machine vision technology according to claim 8, characterized in that, The activation function used is the Sigmoid function, and the activation function used in the convolutional layer is the ReLU function. By introducing non - linear features, the expression ability of the model is increased.

10. The surface anomaly detection algorithm for belt buckles based on machine vision technology according to claim 1, characterized in that, The specific method for determining the detection result is as follows: Calculate the absolute value of the difference between the corresponding pixel values of the original image and the reconstructed image to obtain a difference image; Use the OTSU method to perform binary processing on the difference image, automatically determine a suitable threshold, and convert the difference image into a binary image. Among them, white pixels represent areas where abnormalities may exist, black pixels represent normal areas, pixels with gray values higher than the threshold are set to white, and pixels lower than the threshold are set to black; Use the erosion and dilation techniques in image morphology to process the binary image. Use a 5x5 all - 1 matrix as the structural element to remove noise points and fill holes; Use a contour detection algorithm to find the contours of all white areas in the binary image and calculate the area enclosed by each contour. Find the white area with the largest area and calculate the proportion of its area to the total area of the belt buckle; when this proportion exceeds 10%, it is determined that there is an abnormality in the belt buckle and an alarm message is issued.

Citation Information

Patent Citations

  • Image anomaly detection method based on spatial context variational automatic encoder

    CN114913377A

  • Display module defect detection method based on machine vision

    CN118483233A

  • Railway fastener anomaly detection method and system based on computer vision

    CN119205747A

  • Surface defect detection method and apparatus, and electronic device

    WO2019233166A1