An AI-based 360-degree full-view high-speed defect detection method and system for bottle caps

Through the orthogonal distribution of multiple cameras and AI defect recognition model, the problems of viewing angle limitations and low efficiency of traditional bottle cap inspection are solved, and 360° full-view high-speed inspection of bottle caps is achieved, which improves the comprehensiveness and accuracy of inspection and adapts to the inspection needs of bottle caps of different specifications.

CN120451176BActive Publication Date: 2025-09-05杭州映图智能科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510964963.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-05
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Traditional bottle cap inspection methods have problems such as limited detection viewing angle, low efficiency, weak model generalization ability and insufficient dynamic adaptability, making it difficult to achieve 360° full-view defect detection and high-speed inspection.

Method used

The orthogonal distribution of multiple cameras combined with an endoscope lens is used to obtain full-view images of bottle caps. Through the target surface calibration strategy and image transposition matrix stitching, combined with the AI ​​defect recognition model and high-speed detection mechanism, 360° full-view defect detection of bottle caps can be achieved.

Benefits of technology

It achieves 360° full-view coverage of the inner wall, side wall, and upper and lower end faces of the bottle cap, improves the comprehensiveness and accuracy of detection, adapts to the detection of bottle caps of various specifications, reduces equipment maintenance costs, and ensures the stable operation of high-speed production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451176B_ABST
    Figure CN120451176B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of bottle cap visual inspection, and specifically to an AI-based 360-degree full-viewing angle defect high-speed detection method and system for bottle caps. The method comprises a detection image acquisition step, an image processing and splicing step, a defect identification and detection step, and a detection result verification step. Through a multi-sensor collaborative layout, 360-degree blind-angle detection of bottle caps is achieved, and a high-speed detection mechanism is used to ensure comprehensive detection of bottle caps. A clustering model is constructed using historical defect data, abnormal defects are determined using the Mahalanobis distance, and the model is dynamically updated based on transfer learning to achieve rapid learning of new defects and reduce manual intervention costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bottle cap visual inspection, and specifically to an AI-based 360-degree full-viewing angle defect high-speed detection method and system for bottle caps. Background Art

[0002] In beverage packaging production, plastic bottle caps are critical sealing components, and their quality directly impacts product safety and reliability. Traditional bottle cap defect detection relies primarily on manual visual inspection or fixed-threshold machine vision technology, but these methods suffer from significant shortcomings: Limited inspection field of view: Traditional methods typically only inspect the top surface or a portion of the sidewall of the bottle cap, failing to cover the inner wall, bottom surface, and the entire 360° circumference. This creates blind spots, leading to missed detections of hidden defects such as cracks, burrs, and seal defects. Low inspection efficiency: Manual inspection speeds are limited to only a few dozen per minute and are susceptible to fatigue. While traditional machine vision can increase speed, it often misses or misidentifies bottle caps on high-speed conveyors (e.g., exceeding 1,000 caps per minute) due to image acquisition and processing delays. Weak model generalization: Different bottle cap specifications (diameter, height, and surface texture variations) require individual inspection parameter adjustments, even requiring retraining of the model. This leads to high equipment maintenance costs and makes it difficult to adapt to the flexible production needs of high-variety, small-batch production. Insufficient dynamic adaptability: When the conveyor speed fluctuates or the spacing between bottle caps is uneven, the traditional trigger mechanism can easily cause image acquisition to be misaligned or overlapped, affecting detection accuracy. This problem is particularly prominent under high-speed conditions.

[0003] Therefore, in order to solve the problems existing in the prior art, the present invention proposes an AI-based 360-degree full-viewing high-speed defect detection method and system for bottle caps. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide an AI-based 360-degree full-viewing high-speed defect detection method and system for bottle caps.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] An AI-based high-speed 360-degree full-view defect detection method for bottle caps includes the following steps:

[0007] an inspection image acquisition step, obtaining a full-view image set of the bottle cap to be inspected by using a side wall four-way camera, a top camera, and a bottom endoscope lens, wherein the full-view image set of the bottle cap to be inspected includes a top image, an inner wall image, a first side wall image, a second side wall image, a third side wall image, and a fourth side wall image;

[0008] an image processing and stitching step, wherein a sampling plane corresponding to one image is selected as an optimal target plane from the full-view image set of the bottle caps to be inspected through a target plane calibration strategy, an image transposition matrix is ​​established through a feature matching relationship between each image and the optimal target plane, and the images in the full-view image set of the bottle caps to be inspected are transposed using the image transposition matrix and then stitched together to obtain an image to be inspected;

[0009] a defect recognition and detection step, inputting the image to be detected into a defect recognition model, detecting the bottle cap defects through the defect recognition model, and outputting the detection results;

[0010] The test result verification step verifies the test result by combining the historical test data of the bottle cap with the result verification model and outputs the verification result as accurate or inaccurate. When the test result is inaccurate, the defect recognition model is corrected.

[0011] As a further improvement of the present invention, the target surface calibration strategy includes setting the plane where the image with the most overlapping samples with other perspective images is located as the optimal target surface, setting a number of identification points in the overlapping area of ​​the optimal target surface and other perspective images, and the identification points are marking points of the image features in the corresponding sampling area, performing feature matching on each image with the identification points in the optimal target surface, and establishing an image transposition matrix based on the feature matching results.

[0012] As a further improvement of the present invention, the image transposition matrix includes an affine transformation matrix calculated by least squares method based on identification points of at least three groups of feature matching, the affine transformation matrix includes a rotation matrix, a translation vector and a scaling factor, and according to a preset affine transformation matrix parameter table, each image is adjusted in image direction by the rotation matrix, the image splicing position is adjusted by the translation vector, and the image is scaled according to the scaling factor to obtain an image to be spliced ​​directly with the optimal target surface, and the transposed image is gradually fused from a low-resolution layer to a high-resolution layer by an image pyramid layering technology.

[0013] As a further improvement of the present invention, the defect identification and detection step includes a high-speed detection mechanism, which establishes a bottle cap arrival time prediction model based on the coded signal of the conveyor belt, and calculates the predicted time for the bottle cap to arrive at the detection point on the conveyor belt according to the bottle cap arrival time prediction model. When the predicted time of adjacent bottle caps is less than a preset threshold, it is judged that the distance between the two bottle caps is less than the safe detection distance. At this time, the position of the bottle cap is monitored and detected in real time by the laser tube sensor. When the occlusion rate obtained by comparing the bottle cap sampling image with the historical image is greater than the preset threshold, the previous frame image is called to reconstruct the missing area through the optical flow method to obtain a complete image.

[0014] As a further improvement of the present invention, the construction of the defect recognition model includes establishing a backbone network for multi-layer convolution feature extraction of the inspection image, sequentially generating a low-resolution feature map, a medium-resolution feature map and a high-resolution feature map containing the global contour of the bottle cap, cross-layer connection of the feature maps of different resolutions through a feature pyramid network, fusing the detail information of the high-resolution feature map with the semantic information of the low-resolution feature map, generating a fused feature map with both position accuracy and semantic information, outputting a defect probability heat map of each area of ​​the bottle cap through a positioning convolution layer with a convolution kernel of 1x1, and determining the area in the heat map where the defect probability is greater than a preset threshold through a non-maximum suppression algorithm, generating a defect candidate box and outputting the confidence of the defect type of the feature area corresponding to the defect candidate box through a fully connected layer and a classification function, and when the confidence is greater than the preset confidence threshold, outputting the area as having a defect and outputting the defect type.

[0015] As a further improvement of the present invention, the detection result verification step includes establishing a constraint rule for the defect position through the bottle cap image to be detected. The constraint rule includes that when the damage rate obtained by the defect position detection is greater than a preset damage rate threshold, it is determined to be a suspicious result and the verification result is output as inaccurate, and the current defect feature is input into the historical defect clustering model to obtain historical defect clustering data. When the Mahalanobis distance between the historical defect clustering data and the nearest cluster center is greater than a preset threshold, the defect detection is re-performed.

[0016] As a further improvement of the present invention, the defect recognition model correction includes, when the verification result is inaccurate, adding image samples obtained by enhancing the detection image data into the feature space corresponding to the inaccurate verification result area through transfer learning technology, and performing convolution iterative correction within a preset number of times through the last three convolution layers of the classification sub-model based on the image samples.

[0017] As a further improvement of the present invention, the detection image acquisition step includes that the side wall four-way cameras are evenly distributed along the circumference of the bottle cap in the directions of 0°, 90°, 180°, and 270°, each camera is equipped with a strip light source, and the light source and the side wall four-way camera are synchronously exposed and sampled through a pulse synchronization trigger circuit.

[0018] As a further improvement of the present invention, the step of specification recognition is further included, which obtains the laser coding area of ​​the top image and outputs the diameter and height of the bottle cap to be inspected as specification data in real time through a recognition model constructed by a convolutional neural network and an attention mechanism, and adjusts the sampling distance and focal length of the side wall camera according to the specification data.

[0019] An AI-based, 360-degree, full-view, high-speed bottle cap defect detection system, including:

[0020] The detection image acquisition module acquires a full-view image set of the bottle cap to be inspected through the side wall four-way camera, the top camera and the bottom endoscope lens, wherein the full-view image set of the bottle cap to be inspected includes a top image, an inner wall image, a first side wall image, a second side wall image, a third side wall image and a fourth side wall image;

[0021] An image processing and stitching module selects a sampling plane corresponding to an image in the full-view image set of the bottle cap to be inspected as the optimal target surface through a target surface calibration strategy, establishes an image transposition matrix based on the feature matching relationship between each image and the optimal target surface, and transposes the images in the full-view image set of the bottle cap to be inspected through the image transposition matrix and stitches them together to obtain an image to be inspected;

[0022] a defect recognition and detection module, which inputs the image to be detected into a defect recognition model, detects the bottle cap defects through the defect recognition model, and outputs the detection results;

[0023] The test result verification module verifies the test result by combining the historical test data of the bottle cap with the result verification model and outputs the verification result as accurate or inaccurate. When the test result is inaccurate, the defect recognition model is corrected.

[0024] The beneficial effects of the present invention are:

[0025] (1) By using multiple cameras in orthogonal distribution combined with an endoscope lens, a 360° full-view coverage of the inner wall, side wall, and upper and lower end faces of the bottle cap is achieved, solving the detection blind spot problem of traditional technology, fully capturing details, and improving the comprehensiveness and accuracy of detection.

[0026] (2) The high-speed detection mechanism predicts the arrival time of bottle caps based on the conveyor belt coding signal and monitors the spacing in real time with the laser tube sensor. When the spacing between bottle caps is less than the safe distance, the camera sampling frequency is automatically increased to 500fps, and the image loss in the overlapping area is compensated by the optical flow image reconstruction technology to ensure stable detection of more than 1,000 bottle caps per minute. Dynamic ROI division and multi-threaded processing technology are used to process adjacent bottle cap images in parallel to avoid detection congestion caused by too close spacing and ensure continuous operation of the assembly line.

[0027] (3) By using a general AI model with strong generalization capabilities, it can handle bottle caps of various specifications without the need to train a separate model for each specification. It can automatically learn the surface features of the bottle caps, adapt to the inspection requirements of different batches, simplify equipment configuration, and reduce maintenance and upgrade costs. In addition, it uses historical defect data to build a clustering model, determines abnormal defects through Mahalanobis distance, and dynamically updates the model based on transfer learning to achieve rapid learning of new defects and reduce the cost of manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1This is a flow chart of a high-speed AI-based 360-degree full-view defect detection method for bottle caps of the present invention;

[0029] Figure 2 This is a flowchart of the image processing and splicing steps of the AI-based high-speed 360-degree full-view defect detection method for bottle caps of the present invention;

[0030] Figure 3 This is a defect recognition and detection and high-speed detection mechanism flow chart of the present invention's AI-based 360-degree full-view bottle cap defect high-speed detection method;

[0031] Figure 4 This is a flow chart of the test result verification and model correction of the AI-based high-speed 360-degree full-view defect detection method for bottle caps of the present invention;

[0032] Figure 5 This is a hardware architecture block diagram of the detection system for the AI-based high-speed 360-degree full-view defect detection method for bottle caps of the present invention. DETAILED DESCRIPTION

[0033] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom," "top," "inner," and "outer" refer to directions toward or away from the geometric center of a particular component, respectively.

[0034] An AI-based high-speed 360-degree full-view defect detection method for bottle caps, such as Figures 1 to 5 As shown, the following steps are included:

[0035] an inspection image acquisition step, obtaining a full-view image set of the bottle cap to be inspected by using a side wall four-way camera, a top camera, and a bottom endoscope lens, wherein the full-view image set of the bottle cap to be inspected includes a top image, an inner wall image, a first side wall image, a second side wall image, a third side wall image, and a fourth side wall image;

[0036] In practical applications, four sidewall cameras are installed at 0°, 90°, 180°, and 270° around the bottle cap. Each camera is equipped with a bar light source, and a pulse synchronization trigger circuit achieves synchronized exposure and sampling between the light source and camera. For example, on a beverage bottle cap production line, when the bottle cap moves on the conveyor belt to the inspection area, the four sidewall cameras, the top camera, and the bottom endoscope lens are activated simultaneously to capture images of the bottle cap's top, inner wall, and four sidewalls, forming a full-view image set.

[0037] an image processing and stitching step, wherein a sampling plane corresponding to one image is selected as an optimal target plane from the full-view image set of the bottle caps to be inspected through a target plane calibration strategy, an image transposition matrix is ​​established through a feature matching relationship between each image and the optimal target plane, and the images in the full-view image set of the bottle caps to be inspected are transposed using the image transposition matrix and then stitched together to obtain an image to be inspected;

[0038] According to the target surface calibration strategy, the plane of the image with the most overlapping sampling with other perspective images is selected as the optimal target surface. Under normal circumstances, the top image has the most overlapping sampling areas with other perspective images, so the sampling plane of the top image is set as the optimal target surface. Several identification points are set in the overlapping areas of the top image and other perspective images, such as setting identification points at the edge features of the bottle cap. Feature matching is performed on each image with the identification points in the top image, and based on at least three sets of feature-matched identification points, the affine transformation matrix is ​​calculated by the least squares method to obtain the rotation matrix, translation vector and scaling factor. According to the preset affine transformation matrix parameter table, the direction, position and scale of each image are adjusted to obtain the image to be spliced. Through the image pyramid layering technology, the transposed images are gradually fused from the low-resolution layer to the high-resolution layer to obtain the complete image to be detected.

[0039] a defect recognition and detection step, inputting the image to be detected into a defect recognition model, detecting the bottle cap defects through the defect recognition model, and outputting the detection results;

[0040] The image to be inspected is input into the defect recognition model. The backbone network performs multi-layer convolution feature extraction on the image to generate low-, medium-, and high-resolution feature maps. The feature pyramid network connects feature maps of different resolutions across layers, fuses detail information and semantic information, and generates a fused feature map. The positioning convolution layer outputs a defect probability heat map, and a non-maximum suppression algorithm is used to generate defect candidate boxes. The fully connected layer and classification function output the confidence level of the defect type. When the confidence level is greater than the preset threshold, the area is determined to be defective and the defect type is output. At the same time, the high-speed detection mechanism establishes a bottle cap arrival time prediction model based on the coded signal of the conveyor belt and calculates the predicted time for the bottle cap to arrive at the detection point. When the predicted time of adjacent bottle caps is less than the preset threshold, the laser tube sensor monitors the position of the bottle cap in real time. If the occlusion rate is greater than the preset threshold, the previous frame image is called to reconstruct the missing area using the optical flow method.

[0041] The test result verification step verifies the test result by combining the historical test data of the bottle cap with the result verification model and outputs the verification result as accurate or inaccurate. When the test result is inaccurate, the defect recognition model is corrected.

[0042] Constraints for defect location are established using the bottle cap image to be inspected, such as setting a damage rate threshold of 5%. If the damage rate at a defect location exceeds 5%, the result is considered suspicious, the verification result is output as inaccurate, and the current defect features are input into the historical defect clustering model. If the Mahalanobis distance between the historical defect cluster data and the nearest cluster center exceeds the preset threshold, defect detection is repeated.

[0043] Specifically, such as Figures 1 to 5 As shown, the target surface calibration strategy includes setting the plane where the image with the most overlapping samples with other view images is located as the optimal target surface, setting a number of identification points in the overlapping area between the optimal target surface and other view images, and the identification points are marking points of the image features in the corresponding sampling area, performing feature matching on each image and the identification points in the optimal target surface, and establishing an image transposition matrix based on the feature matching results.

[0044] In the image processing and stitching steps, Figure 2 As shown in Figure 2, the implementation of the target surface calibration strategy includes:

[0045] The optimal target surface is selected by comparing the overlapping sampling areas of each viewpoint image (top, inner wall, and four-way sidewall) through image preprocessing (such as grayscale and edge detection). The plane containing the image with the largest overlap area with the other viewpoint images is selected as the optimal target surface. For example, if the inner wall image captured by the bottom endoscope lens and the four-way sidewall images all have a circular overlap area, the inner wall image plane is preferentially selected as the optimal target surface because it covers the key inspection area on the inside of the bottle cap. When the overlap rate of a certain viewpoint image is higher than that of other viewpoints by more than 20%, that is, the overlap rate is ≥20%, it is determined to be the optimal target surface. The setting of the marker points and feature matching are carried out. Within the overlapping area between the optimal target surface and the other viewpoint images, at least five marker points (such as the inflection point of the bottle cap edge and the color mutation point) are automatically or manually selected. Each marker point corresponds to a unique image feature (such as SIFT feature or ORB feature) within the sampling area. The FLANN fast approximate nearest neighbor algorithm is used to perform feature matching between the marker points of each viewpoint image and the optimal target surface, and a cross-viewpoint correspondence relationship between the marker points is established. For example, the marker point A of the inner wall image and the marker point A' of the 0° side wall image are matched through the cosine similarity of the feature vectors (the threshold is set to 0.9) to ensure the consistency of the geometric positions.

[0046] Specifically, such as Figures 1 to 5As shown, the image transposition matrix includes an affine transformation matrix obtained by least squares calculation based on identification points of at least three sets of feature matching, the affine transformation matrix includes a rotation matrix, a translation vector and a scaling factor, and the image direction of each image is adjusted by the rotation matrix according to a preset affine transformation matrix parameter table. The image splicing position is adjusted by the translation vector and the image is scaled according to the scaling factor to obtain an image to be spliced ​​directly with the optimal target surface, and the image to be spliced ​​is gradually fused with the transposed image from the low-resolution layer to the high-resolution layer through the image pyramid layering technology.

[0047] Based on at least three sets of successfully matched identification point pairs, the affine transformation matrix is ​​solved by the least squares method:

[0048] ,satisfy: , where (x, y) is the coordinate of the optimal target surface marker point, (x', y') is the coordinate of the corresponding marker point of the image to be transposed, and a, b, c, d, e, and f represent the linear transformation parameters, respectively. a represents the combined factor of the x-direction scaling factor and the y-direction shearing factor, and is related to the cosine value of the rotation angle; b represents the combined factor of the x-direction shearing factor and the y-direction scaling factor, and is related to the sine value of the rotation angle; d represents the combined factor of the y-direction shearing factor and the x-direction scaling factor, and is related to the negative sine value of the rotation angle; e represents the combined factor of the y-direction scaling factor and the x-direction shearing factor, and is related to the cosine value of the rotation angle; c represents the image translation in the x-axis direction; and f represents the image translation in the y-axis direction. The matrix parameters are iteratively optimized to make the matching error (root mean square error) less than 1 pixel, ensuring the accuracy of the geometric transformation. The above equations are solved by the least squares method to make the matching error of all marker point pairs satisfy: <1 pixel, thus ensuring that the image coordinate error after geometric transformation is controlled at the sub-pixel level, enhancing accuracy.

[0049] Image transposition and stitching processing:

[0050] Orientation: Setting the rotation matrix components Rotate the image counterclockwise (such as 90°, 180°) to make the coordinate system of the image to be stitched consistent with the coordinate system of the optimal target surface. R is the rotation and scaling matrix, and its determinant value is , indicating the scaling ratio, when the rotation state is ideal .

[0051] Position and scale adjustment: Align the image center to the stitching origin of the optimal target surface by translating the vector (c, f), and adjust the image according to the scaling factor. or Scale the image proportionally to avoid distortion.

[0052] Image pyramid layered fusion: The image to be stitched is constructed into a three-layer pyramid (low-resolution layer: 1 / 4 original image size, medium-resolution layer: 1 / 2, high-resolution layer: original image). Starting from the low-resolution layer, the Laplacian pyramid algorithm is used to fuse the layers layer by layer. Gaussian blur and interpolation calculation are used to retain details at each scale, ultimately generating a seamless image to be detected.

[0053] Specifically, such as Figure 3 As shown, the defect recognition and detection step includes a high-speed detection mechanism, which establishes a bottle cap arrival time prediction model based on the coded signal of the conveyor belt, and calculates the predicted time when the bottle caps on the conveyor belt arrive at the detection point according to the bottle cap arrival time prediction model. When the predicted time of adjacent bottle caps is less than a preset threshold, it is judged that the distance between the two bottle caps is less than the safe detection distance. At this time, the position of the bottle caps is monitored and detected in real time by the laser tube sensor. When the occlusion rate obtained by comparing the bottle cap sampling image with the historical image is greater than the preset threshold, the previous frame image is called to reconstruct the missing area through the optical flow method to obtain a complete image.

[0054] The bottle cap arrival time prediction model installs an incremental encoder on the conveyor drive shaft to collect pulse signals in real time to calculate the linear speed. , v represents the linear velocity of the conveyor belt drive shaft, Pc represents the number of pulses, Pe represents the pulse equivalent, and is the conveyor belt displacement corresponding to each pulse. Indicates the sampling time interval. When a bottle cap passes the detection point, its arrival time t is recorded. i , and predict the arrival time of the next bottle cap based on the conveyor belt speed v , where L i,i+1The actual spacing between adjacent bottle caps on the conveyor belt is initially preset using historical inspection data and adjusted in real time based on the previous inspection data during the current inspection process. When the predicted time difference between adjacent bottle caps is less than the preset time threshold, it indicates that the bottle caps are too close together, and the laser alignment sensor is triggered. The laser alignment sensor (with an accuracy of 0.1mm) is installed 10cm upstream of the inspection point. When the previous bottle cap is detected, the timing is started. If the next bottle cap is detected blocking the laser beam within the time threshold, the actual spacing L is determined to be less than the distance the conveyor belt moves during this time. At this time, high-speed inspection is initiated. When the spacing is determined to be too close, the camera sampling frequency is increased from the default 200fps to 500fps to ensure that multiple bottle cap images are captured in a short period of time. At this time, dynamic ROI division is triggered to mark the overlapping areas of adjacent bottle caps to avoid missed detections. By comparing the current frame with the previous frame, the occlusion area is identified through pixel XOR operation. The occlusion rate is the ratio of the number of XOR non-zero pixels to the total number of pixels of the current bottle cap area. When the occlusion rate is greater than the preset occlusion rate threshold, the corresponding area of ​​the previous frame image is called and the missing part of the image is reconstructed using the bidirectional optical flow method. The image reconstruction steps of the bidirectional optical flow method include:

[0055] The bidirectional optical flow field estimation step uses the PWC-Net network and inputs the current frame image I t and the previous complete image I t-1 , calculated from I at multiple levels of the image pyramid t to I t-1 The optical flow field u t→t-1 , represents the corresponding position of the current frame pixel in the previous frame, and outputs the forward optical flow map u f , each pixel (m, n) corresponds to a vector (u m ,u n ), which represents the coordinate of the pixel in the previous frame (m+u m , n+u n ); Backward optical flow calculation, input the previous frame image I t-1 and the current frame image I t , using the same optical flow model to calculate from I t-1 to I t The optical flow field u t-1→t , get the backward optical flow map u b , At the same time, the forward optical flow and the backward optical flow meet the consistency, (m, n) = (m+u m , n+u n )+u b (m+u m , n+u n), the pixels that do not meet the constraint are marked as unreliable optical flow points, and interpolation is performed to repair the pixels; in the occlusion area positioning step, the previous frame pixel (m',n') is projected to the current frame through the forward optical flow to obtain the projection coordinates (m0,n0) = (m'+u f (m',n')), calculate the difference between the projected pixel and the pixel value at the corresponding position in the current frame e(m',n')=||I t (x, y)-I t-1 (m',n')||, when the difference value e is significantly greater than the preset threshold (usually set to 20 grayscale values), the area is determined to be the occlusion area, that is, the bottle cap area newly added in the current frame has no corresponding pixel in the previous frame. For the pixels (m,n) in the occlusion area of ​​the current frame, the backward optical flow (u b (m,n) reverses its position in the previous frame ((m',n')=(mu b (m,n)). If (m',n') exceeds the boundary of the previous frame image or the corresponding pixel is an occluded area, it is determined to be a valid missing area, and image reconstruction is performed at this time; the missing area boundary interpolation step extracts the contour boundary of the missing area of ​​the current frame, and for each pixel (m,n) on the boundary, find its corresponding position in the previous frame (m',n')=(mu f (m,n)), using the optical flow vectors of the four reliable pixels around the boundary pixel, the weighted average is used to obtain the optical flow estimate of the pixels inside the missing area: ,in The pixel value of the previous frame is propagated along the direction of the optical flow, and the transition is smoothed by the Gaussian blur kernel (kernel size 3×3, standard deviation 1.5) to avoid block effects in the reconstructed area. In the spatiotemporal information fusion reconstruction step, according to the forward optical flow u f , the pixel value I of the non-occluded area of ​​the previous frame t-1 (m',n') is mapped to the missing area of ​​the current frame: , where (m, n) are the pixel coordinates of the missing region. Temporal filtering is performed on the reconstructed region of three consecutive frames, suppressing inter-frame noise through weighted averaging to improve the temporal consistency of the reconstructed image. The SSIM value of the reconstructed region is calculated compared to adjacent valid regions. If the SSIM value is less than 0.7, the reconstruction quality is considered insufficient, triggering the following corrections: The interpolation range is expanded to include pixels from the previous frame that are farther from the boundary in the reconstruction; a generative adversarial network (GAN) (such as CycleGAN) is used to perform texture restoration on the reconstructed region to improve detail fidelity. The reconstructed image is used as the new current frame, and bidirectional optical flow is recalculated. Unreliable areas in the optical flow field are updated, forming a closed-loop correction mechanism to ensure reconstruction accuracy in subsequent frames.

[0056] Specifically, such as Figures 1 to 5As shown, the construction of the defect recognition model includes establishing a backbone network for multi-layer convolution feature extraction of the inspection image, sequentially generating a low-resolution feature map, a medium-resolution feature map and a high-resolution feature map containing the global contour of the bottle cap, performing cross-layer connection on the feature maps of different resolutions through a feature pyramid network, fusing the detail information of the high-resolution feature map with the semantic information of the low-resolution feature map, generating a fused feature map with both position accuracy and semantic information, outputting a defect probability heat map of each area of ​​the bottle cap through a positioning convolution layer with a convolution kernel of 1x1, and determining the area in the heat map where the probability of a defect is greater than a preset threshold through a non-maximum suppression algorithm, generating a defect candidate box and outputting the confidence of the defect type of the feature area corresponding to the defect candidate box through a fully connected layer and a classification function. When the confidence is greater than the preset confidence threshold, the area is output as having a defect and the defect type is output.

[0057] The backbone network design uses ResNet-50 as the feature extraction backbone, which contains 50 convolutional layers and is divided into two stages:

[0058] Stage 1: A 7×7 convolutional layer (stride 2) extracts the global contour and outputs a low-resolution feature map (size 1 / 4 of the original image, number of channels 64).

[0059] Stage 2: Use the residual block to gradually reduce the feature map size (1 / 8, 1 / 16, 1 / 32), increase the number of channels (128 → 256 → 512), and generate medium-resolution and high-resolution feature maps.

[0060] Feature Pyramid Network (FPN) fusion restores the high-resolution feature map (such as the output of stage 5, size 1 / 32) to the stage 2 size (1 / 8) through upsampling (nearest neighbor interpolation), and adds it element-by-element with the medium-resolution feature map of stage 2, fusing detail information (such as crack edges) and semantic information (such as the "crack" category), ultimately generating three layers of fused feature maps (sizes 1 / 8, 1 / 16, and 1 / 32), corresponding to different detection scales.

[0061] To generate a defect probability heatmap, the fused feature map is passed through a 1×1 convolutional layer (number of channels = number of defect categories + 1) to output a heatmap. Each pixel value represents the probability that the corresponding location belongs to a certain defect type (for example, 0.8 indicates an 80% probability of being a crack). The non-maximum suppression (NMS) algorithm is used with an intersection-over-union (IoU) threshold of 0.5 to retain the defect candidate boxes with the highest probability.

[0062] For classification and positioning output, the feature area corresponding to the defect candidate box is processed through a fully connected layer (number of neurons = number of defect categories) and a Softmax function, outputting the confidence level for each category. When the confidence level is greater than 0.9, it is considered a valid defect and the coordinates, category, and probability are output.

[0063] Specifically, such as Figures 1 to 5 As shown, the detection result verification step includes establishing a constraint rule for the defect position through the bottle cap image to be detected. The constraint rule includes that when the damage rate obtained by the defect position detection is greater than a preset damage rate threshold, it is determined to be a suspicious result and the verification result is output as inaccurate, and the current defect feature is input into the historical defect clustering model to obtain historical defect clustering data. When the Mahalanobis distance between the historical defect clustering data and the nearest cluster center is greater than a preset threshold, the defect detection is re-performed.

[0064] Defect location constraint rules are established, with a preset damage rate threshold of 20% for key functional areas of the bottle cap (such as the inner wall sealing surface and the anti-theft ring connection point). For example, if the pixel area of ​​the inner wall sealing surface with defects accounts for 25% of the total area of ​​the area, a suspicious result is triggered. The area is pre-annotated using a semantic segmentation model (such as U-Net), and masks are generated for each functional area for rapid screening of defect locations.

[0065] The historical defect clustering model uses the DBSCAN density clustering algorithm to cluster historical defect features (such as location, shape, and color histogram) and automatically determines the number of cluster centers. It calculates the Mahalanobis distance between the current defect feature and the nearest cluster center. ,in is the cluster center mean vector, is the covariance matrix. M When the value is >2.5, it is determined to be an abnormal defect and re-inspection is triggered.

[0066] Specifically, such as Figure 4 As shown, the defect recognition model correction includes, when the verification result is inaccurate, adding image samples obtained by enhancing the detection image data in the feature space corresponding to the inaccurate verification result area through transfer learning technology, and performing convolution iterative correction within a preset number of times through the last three convolution layers of the classification sub-model based on the image samples.

[0067] When the verification result is inaccurate, the model correction process is as follows:

[0068] Data enhancement and feature space expansion: Data enhancement is performed on images in areas with inaccurate verification results, including: geometric transformation: rotation (±10°), translation (±5 pixels), scaling (0.9-1.1 times); color transformation: brightness adjustment (±20%), contrast adjustment (±15%), and adding Gaussian noise (mean 0, variance 0.01); 200-500 enhanced samples are generated to expand the diversity of defect features.

[0069] Transfer learning and convolution iterative correction: freeze the first 10 convolutional layers of the defect recognition model and fine-tune only the last three layers (such as layer 4, layer 5 and positioning convolution layer of ResNet). Use the stochastic gradient descent (SGD) algorithm, set the learning rate to 0.001, iterate 10 times, and optimize the loss function (cross entropy loss + smooth L1 loss) to enhance the model's response to newly added defect features.

[0070] Specifically, such as Figures 1 to 5 As shown, the detection image acquisition step includes: the side wall four-way cameras are evenly distributed along the circumference of the bottle cap in the directions of 0°, 90°, 180°, and 270°, each camera is equipped with a strip light source, and a pulse synchronization trigger circuit is used to synchronize the light source and the side wall four-way cameras for exposure and sampling.

[0071] Camera and light source layout:

[0072] Four-way side wall camera: Four industrial area array cameras (resolution 2048×2048, frame rate 200fps) are evenly distributed along the 0°, 90°, 180°, and 270° directions, with the optical axis perpendicular to the bottle cap side wall and 10-15cm away from the bottle cap surface; strip light source: Each camera is equipped with an independent LED strip light source (525nm wavelength green light), which illuminates the bottle cap side wall at a 45° angle to enhance the surface texture contrast and suppress reflections; pulse synchronization trigger circuit: The FPGA generates a synchronous trigger signal to ensure that the light source exposure time (100μs) is strictly aligned with the camera sampling time to avoid motion blur.

[0073] Bottom endoscope lens: It uses an 8mm diameter telecentric endoscope lens with a built-in ring light source. It penetrates 5-8mm into the bottle cap and collects 360° images of the inner wall.

[0074] Top camera: 1 industrial camera (resolution 1280×1024) shoots the top surface of the bottle cap vertically downward for laser code recognition and top surface defect detection.

[0075] Specifically, such as Figures 1 to 5 As shown, the system also includes a specification recognition step, which obtains the laser coding area of ​​the top image and outputs the diameter and height of the bottle cap to be inspected as specification data in real time through a recognition model constructed by a convolutional neural network and an attention mechanism, and adjusts the sampling distance and focal length of the side wall camera according to the specification data.

[0076] To locate the laser coding area, the Hough circle detection algorithm is used to locate the circular area on the bottle cap's top surface. This is combined with OCR text recognition to extract the laser coding specifications (e.g., "D32×H15" represents a diameter of 32mm and a height of 15mm). The recognition model uses a convolutional neural network and an attention mechanism. The SE attention module is added after the convolutional layer. This module adjusts channel weights to focus on the coded character area, improving small object recognition accuracy. The model input is a 64×64 ROI image of the coded area, and the output is a numerical regression of the diameter and height, with an accuracy of ±0.1mm.

[0077] The camera parameters are adaptively adjusted. According to the specifications, the camera bracket is driven by a stepper motor: for every 1mm increase in diameter, the sampling distance of the side wall camera increases by 0.5cm, and the focal length is adjusted synchronously; when the height changes, the top camera automatically adjusts the focal length through the electric zoom lens to ensure clear imaging of the coding area.

[0078] An AI-based 360-degree full-view high-speed defect detection system for bottle caps, such as Figure 5 Shown, including:

[0079] The detection image acquisition module acquires a full-view image set of the bottle cap to be inspected through the side wall four-way camera, the top camera and the bottom endoscope lens, wherein the full-view image set of the bottle cap to be inspected includes a top image, an inner wall image, a first side wall image, a second side wall image, a third side wall image and a fourth side wall image;

[0080] An image processing and stitching module selects a sampling plane corresponding to an image in the full-view image set of the bottle cap to be inspected as the optimal target surface through a target surface calibration strategy, establishes an image transposition matrix based on the feature matching relationship between each image and the optimal target surface, and transposes the images in the full-view image set of the bottle cap to be inspected through the image transposition matrix and stitches them together to obtain an image to be inspected;

[0081] a defect recognition and detection module, which inputs the image to be detected into a defect recognition model, detects the bottle cap defects through the defect recognition model, and outputs the detection results;

[0082] The test result verification module verifies the test result by combining the historical test data of the bottle cap with the result verification model and outputs the verification result as accurate or inaccurate. When the test result is inaccurate, the defect recognition model is corrected.

[0083] The above shows and describes the basic features, principles, and advantages of the present invention. It should be noted that the present invention is not limited to the above embodiments, which are only some embodiments. Without departing from the spirit and scope of the present invention, various improvements and supplements made are considered to be within the scope of protection of the present invention.

Claims

1. An AI-based high-speed 360-degree full-view defect detection method for bottle caps, characterized in that: The steps include: an inspection image acquisition step, obtaining a full-view image set of the bottle cap to be inspected by using a side wall four-way camera, a top camera, and a bottom endoscope lens, wherein the full-view image set of the bottle cap to be inspected includes a top image, an inner wall image, a first side wall image, a second side wall image, a third side wall image, and a fourth side wall image; an image processing and stitching step, wherein a sampling plane corresponding to one image is selected as an optimal target plane from the full-view image set of the bottle caps to be inspected through a target plane calibration strategy, an image transposition matrix is ​​established through a feature matching relationship between each image and the optimal target plane, and the images in the full-view image set of the bottle caps to be inspected are transposed using the image transposition matrix and then stitched together to obtain an image to be inspected; a defect recognition and detection step, inputting the image to be detected into a defect recognition model, detecting the bottle cap defects through the defect recognition model, and outputting the detection results; The test result verification step uses the result verification model to verify the test result by combining the historical test data of the bottle cap and outputs the verification result as accurate or inaccurate. When the test result is inaccurate, the defect recognition model is corrected; The target surface calibration strategy includes setting the plane containing the image with the most overlapping samples with other viewpoint images as the optimal target surface, setting a number of identification points in the overlapping area of ​​the optimal target surface with other viewpoint images, wherein the identification points are marking points corresponding to image features in the sampling area, performing feature matching between each image and the identification points in the optimal target surface, and establishing an image transposition matrix based on the feature matching results; The image transposition matrix includes an affine transformation matrix obtained by least squares calculation based on identification points of at least three sets of feature matching, the affine transformation matrix including a rotation matrix, a translation vector and a scaling factor, adjusting the image direction of each image using the rotation matrix according to a preset affine transformation matrix parameter table, adjusting the image splicing position using the translation vector and scaling the image according to the scaling factor to obtain an image to be spliced ​​directly with the optimal target surface, and gradually fusing the transposed image from a low-resolution layer to a high-resolution layer using an image pyramid layering technology; The construction of the defect recognition model includes establishing a backbone network for multi-layer convolution feature extraction of the image to be detected, sequentially generating a low-resolution feature map, a medium-resolution feature map, and a high-resolution feature map containing the global contour of the bottle cap, cross-layer connection of the feature maps of different resolutions through a feature pyramid network, fusing the detail information of the high-resolution feature map with the semantic information of the low-resolution feature map, generating a fused feature map with both position accuracy and semantic information, outputting a defect probability heat map of each area of ​​the bottle cap through a positioning convolution layer with a convolution kernel of 1x1, and determining through a non-maximum suppression algorithm that there is an area in the heat map with a defect probability greater than a preset threshold, generating a defect candidate box, and outputting the confidence of the defect type of the feature area corresponding to the defect candidate box through a fully connected layer and a classification function. When the confidence is greater than the preset confidence threshold, the area is output as having a defect and the defect type is output.

2. The AI-based high-speed 360-degree full-view defect detection method for bottle caps according to claim 1, characterized in that: The defect identification and detection step includes a high-speed detection mechanism, which establishes a bottle cap arrival time prediction model based on the coded signal of the conveyor belt, and calculates the predicted time when the bottle caps on the conveyor belt arrive at the detection point according to the bottle cap arrival time prediction model. When the predicted time of adjacent bottle caps is less than a preset threshold, it is judged that the distance between the two bottle caps is less than the safe detection distance. At this time, the position of the bottle caps is monitored and detected in real time by the laser tube sensor. When the occlusion rate obtained by comparing the bottle cap sampling image with the historical image is greater than the preset threshold, the previous frame image is called to reconstruct the missing area through the optical flow method to obtain a complete image.

3. The AI-based high-speed 360-degree full-view defect detection method for bottle caps according to claim 1, characterized in that: The detection result verification step includes establishing a constraint rule for the defect position based on the bottle cap image to be inspected. The constraint rule includes: when the damage rate obtained by the defect position detection is greater than a preset damage rate threshold, it is determined to be a suspicious result and the verification result is output as inaccurate; the current defect feature is input into the historical defect clustering model to obtain historical defect clustering data; when the Mahalanobis distance between the historical defect clustering data and the nearest cluster center is greater than a preset threshold, the defect detection is re-performed.

4. The AI-based high-speed 360-degree full-view defect detection method for bottle caps according to claim 1, characterized in that: The defect recognition model correction includes, when the verification result is inaccurate, adding image samples obtained by enhancing the image data to be detected in the feature space corresponding to the area of ​​the inaccurate verification result through transfer learning technology, and performing convolution iterative correction within a preset number of times through the last three convolution layers of the classification sub-model based on the image samples.

5. The AI-based high-speed 360-degree full-view defect detection method for bottle caps according to claim 1, characterized in that: The detection image acquisition step includes: the four-way side wall cameras are evenly distributed along the circumference of the bottle cap at 0°, 90°, 180°, and 270° directions; each camera is equipped with a strip light source; and a pulse synchronization trigger circuit is used to synchronize the light source and the four-way side wall cameras for exposure and sampling.

6. The AI-based high-speed 360-degree full-view defect detection method for bottle caps according to claim 1, characterized in that: The system also includes a specification recognition step, which obtains the laser coding area of ​​the top image and outputs the diameter and height of the bottle cap to be inspected as specification data in real time through a recognition model constructed by a convolutional neural network and an attention mechanism, and adjusts the sampling distance and focal length of the side wall four-way camera according to the specification data.

7. An AI-based 360-degree full-view high-speed defect detection system for bottle caps, applicable to the AI-based 360-degree full-view high-speed defect detection method for bottle caps according to any one of claims 1 to 6, characterized in that: include: The detection image acquisition module acquires a full-view image set of the bottle cap to be inspected through the side wall four-way camera, the top camera and the bottom endoscope lens, wherein the full-view image set of the bottle cap to be inspected includes a top image, an inner wall image, a first side wall image, a second side wall image, a third side wall image and a fourth side wall image; An image processing and stitching module selects a sampling plane corresponding to an image in the full-view image set of the bottle cap to be inspected as the optimal target surface through a target surface calibration strategy, establishes an image transposition matrix based on the feature matching relationship between each image and the optimal target surface, and transposes the images in the full-view image set of the bottle cap to be inspected through the image transposition matrix and stitches them together to obtain an image to be inspected; a defect recognition and detection module, which inputs the image to be detected into a defect recognition model, detects the bottle cap defects through the defect recognition model, and outputs the detection results; The test result verification module verifies the test result by combining the historical test data of the bottle cap with the result verification model and outputs the verification result as accurate or inaccurate. When the test result is inaccurate, the defect recognition model is corrected.

Citation Information

Patent Citations

  • Visual identification method and device based on computer processing

    CN119180988A

  • Offshore wind power blade defect detection method

    CN119722632A