Method and apparatus for detecting foreign object inside glass bottle and storage medium
Patent Information
- Application Number
- PCT/CN2025/103754
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2025-06-26
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025103754_01102026_PF_FP_ABST
Abstract
Description
Methods, devices and storage media for detecting foreign objects inside glass bottles Technical Field
[0001] This invention relates to the field of foreign object detection in packaged food, and in particular to a method, apparatus and storage medium for detecting foreign objects inside glass bottles. Background Technology
[0002] With the continuous development of industrial automation, quality inspection technology for bottled products has become an important part of the food, pharmaceutical and other industries. As a common packaging material, the detection of foreign objects inside glass bottles is of great significance to ensuring product quality and consumer safety.
[0003] While some existing technologies offer solutions for detecting defects in glass bottles—for example, Chinese patent CN117351314A discloses a method and system for identifying glass bottle defects—the light source used in this method is a visible light source. This light source can only be used to detect defects in glass bottles with high light transmittance. When the glass bottle contains related products or has labels affixed to it, the light cannot penetrate the glass to form an image, thus it cannot detect foreign objects inside the bottle. Furthermore, Chinese patent CN118090743A provides a multimodal quality inspection system for porcelain wine bottles; similarly, its purpose is also to detect defects in the bottle itself, and therefore it also relies on visible light images.
[0004] In response, some improved existing technologies enhance penetration by replacing the light source with an X-ray source, enabling detection inside thicker glass bottles. For example, Chinese patent CN112508930A discloses a deep learning-based method and device for detecting foreign objects in food, and Chinese patent CN 113567478A discloses a novel single-source, three-view X-ray foreign object detection system. Both use X-ray sources to detect foreign objects inside food packaging, with CN 113567478A specifically targeting foreign object detection inside glass jars. However, this approach still suffers from the following problems: due to the complex structural characteristics of glass bottles, such as varying thicknesses and densities in the body, cap, and bottom, traditional single-voltage X-ray imaging technology struggles to effectively detect foreign objects in different areas within the same image. Existing X-ray detection methods typically rely on a single voltage or simple image processing algorithms for foreign object identification. While this approach may perform well in detecting foreign objects in thinner areas of the bottle, insufficient image contrast and resolution in denser areas like the bottom or cap can lead to decreased detection accuracy. Furthermore, excessively increasing the voltage may cause overexposure in the bottle area, thus obscuring minute foreign objects. Therefore, obtaining clear images in areas of different densities and accurately detecting foreign objects has become a pressing technical challenge.
[0005] In response, although Chinese patent CN 113567478 A attempted to use multiple perspectives for shooting, and the different travel paths corresponding to different perspectives could reduce blind spots, and achieved some success in detecting foreign objects at the bottom of the bottle, it still has the following shortcomings:
[0006] 1. Some foreign matter suspended in the solution cannot be effectively detected. Therefore, on the one hand, the types of foreign matter that can be detected are limited, and on the other hand, the product needs to be allowed to stand for a long time after filling before foreign matter detection can be carried out, which undoubtedly reduces the turnover efficiency of the entire production line.
[0007] 2. The deviation of the three views used is large, making image registration difficult. If forced fusion is performed, it may mislead the model and cause the model to fail to converge or specialize. Therefore, images from different perspectives can only be preprocessed and feature extracted separately before classification, which undoubtedly reduces efficiency. Summary of the Invention
[0008] The purpose of this invention is to solve the problem in existing technologies that require a long period of settling after filling before foreign object detection can be performed. This invention provides a method, device, and storage medium for detecting foreign objects inside glass bottles. It uses X-ray sources with different voltages at the same viewing angle, resulting in all single-channel X-ray images having the same viewing angle. This simplifies registration and avoids new errors introduced by scaling and fusion. Furthermore, the use of different voltages effectively covers imaging requirements for regions of different densities. The output image, obtained through high-low frequency separation, high-frequency enhancement, and subsequent stitching and convolution, clearly displays foreign objects suspended in the glass bottle. This improves detection accuracy and reliability while allowing the foreign object detection process to be placed close to the filling process, eliminating the need for post-filling settling and significantly improving the overall production line efficiency.
[0009] The objective of this invention can be achieved through the following technical solutions:
[0010] A method for detecting foreign objects inside a glass bottle includes:
[0011] Step S1: For the glass bottle to be inspected, acquire multiple X-ray images of different voltages from the same viewing angle;
[0012] Step S2: Align all the acquired X-ray images and synthesize them into a multi-channel first image;
[0013] Step S3: After convolution and upsampling the first image to obtain the low-frequency feature map, subtract the low-frequency feature map from the first image to obtain the high-frequency feature map;
[0014] Step S4: Enhance the high-frequency feature map;
[0015] Step S5: After fusing the enhanced high-frequency feature map and the upsampled low-frequency feature map, perform convolution to obtain the output image;
[0016] Step S6: Input the output image into the target detection model to obtain the foreign object detection result.
[0017] In step S1, three X-ray images with different voltages are acquired from the same viewpoint, and the first image is a three-channel image.
[0018] The first convolution kernel in step S3 has a size of 3*3 and a stride of 2.
[0019] The upsampling process in step S3 uses bilinear interpolation, and the size of the low-frequency feature map is the same as the size of the first image.
[0020] In the enhancement process of step S4, the enhanced high-frequency feature map is: F EP =Concat(F H *Conv4,(F H *Conv2*CBAM*Conv3))
[0021] Wherein: F H For the high-frequency feature map before enhancement, F EP For the enhanced high-frequency feature map, Conv2 is the second convolution kernel, Conv3 is the third convolution kernel, Conv4 is the fourth convolution kernel, CBAM is the attention module operator, and Concat is the concatenation operation.
[0022] Step S5 includes:
[0023] Step S5-1: The enhanced high-frequency feature map and the upsampled low-frequency feature map are stitched together along the channel dimension to obtain the first intermediate image;
[0024] Step S5-2: Perform a convolution operation on the first intermediate image using the fifth convolution kernel to obtain the output image.
[0025] All convolutional kernels and operators in steps S2 to S5 are obtained through deep learning training. The training process is adjusted based on a first loss function, which is:
[0026] Where: L is the first loss function. T(H) represents the normalized pixel value of the pixel in the i-th row and j-th column of the output image obtained in step S5. i,j Let be the normalized pixel value of the pixel in the i-th row and j-th column of the output image from the training samples, and γ be the dynamic weight amplitude control parameter. It is an L1 norm.
[0027] The normalized pixel value for any pixel is:
[0028] in: Let be the value of the k-th channel in the normalized pixel value of the pixel in the i-th row and j-th column. Let be the value of the k-th channel among the pixel values of the i-th row and j-th column before normalization. Let be the minimum value of all channels of the pixel in the i-th row and j-th column before normalization. It represents the maximum value of all channels of the pixel in the i-th row and j-th column before normalization.
[0029] A foreign object detection device for glass bottles includes a memory, a processor, and a program stored in the memory, wherein the processor executes the program to implement the method described above.
[0030] A storage medium having a program stored thereon, which, when executed, implements the method described above.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] 1. By using X-ray sources with different voltages but the same viewing angle, all single-channel X-ray images obtained have the same viewing angle, making registration easier and avoiding new errors introduced by scaling and fusion. In addition, the use of different voltages can effectively cover the imaging needs of different density areas. The output image obtained by sequentially separating high and low frequencies, enhancing high frequencies, and then stitching and convolving can clearly display foreign objects suspended in glass bottles. While improving detection accuracy and reliability, the foreign object detection process can be placed close to the filling process, eliminating the need for post-filling settling and greatly improving the overall production line efficiency.
[0033] 2. Compared to conventional convolution and pooling enhancement methods, this method uses three convolutions and one attention layer to enhance high-frequency features, highlighting their details and clarity while completely ignoring contours and background. The enhanced high-frequency feature map and the upsampled low-frequency feature map are then concatenated and re-convolved to compensate for the lack of contours in the high-frequency feature map, thus enabling the identification of foreign objects suspended in various positions in the glass bottle.
[0034] 3. The first loss function used can assign higher weights to pixels with larger prediction errors, thereby better handling areas with inaccurate predictions. At the same time, the normalization of each pixel based on the three-channel linear mapping can fully integrate the information from different channels, so that the details and contours of each voltage imaging can be fully displayed. Attached Figure Description
[0035] Figure 1 is a schematic diagram of the main steps of the method of the present invention;
[0036] Figure 2 is a schematic diagram of an X-ray image under a single voltage. Detailed Implementation
[0037] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0038] Example 1
[0039] A method for detecting foreign objects inside a glass bottle, as shown in Figure 1, includes:
[0040] Step S1: For the glass bottle to be inspected, acquire multiple X-ray images of different voltages from the same viewing angle;
[0041] Unlike existing foreign object detection methods, such as CN 113567478 A, this application employs a single-view approach. The single-view method offers advantages such as low material costs, simple on-site setup, and convenient maintenance. This application uses three different voltages to photograph a glass bottle from the same viewpoint. Due to the very short shooting time and normal conveyor belt speed, the deviation between the three X-ray images is minimal. The specific shooting interval depends on the performance of the electronic control components on-site; the CCD processing speed is not a limiting factor. Generally, the faster the switching speed supported by the X-ray source and its driver, the smaller the interval can be, resulting in a smaller positional deviation between the three X-ray images. A deviation less than the physical length corresponding to one pixel is generally considered negligible.
[0042] Specifically, in this embodiment, three X-ray images with different voltages are acquired from the same viewpoint, as shown in Figure 2. Each X-ray image is a single-channel image taken under a single voltage. After being combined, a three-channel first image can be obtained.
[0043] Because different voltages are used, the energy of the light emitted by the X-ray source varies under different voltages, which can effectively cover the imaging needs of different density areas.
[0044] Step S2: Align all the acquired X-ray images and synthesize them into a multi-channel first image. Generally, since the shooting angles are the same, if the deviation is less than the physical length corresponding to one pixel, the single image is aligned by default. Of course, in some embodiments, if the deviation is too large, alignment can be achieved by translation.
[0045] Furthermore, in this embodiment, the alignment process also includes cropping, cropping out the glass bottle portion as much as possible while maintaining a smaller overall image size. This effectively reduces the computational requirements in subsequent processing and also reduces noise. During this cropping process, the specific cropping area can be determined using a medium-voltage imaging image. In a medium-voltage imaging image, the distinction between the background and foreground is more pronounced, improving cropping accuracy. Of course, other cropping methods can also be used in other embodiments.
[0046] Step S3: After convolution and upsampling the first image to obtain the low-frequency feature map, subtract the low-frequency feature map from the first image to obtain the high-frequency feature map;
[0047] In this embodiment, the size of the first convolution kernel in the convolution process is 3*3, the stride is 2, the upsampling process uses bilinear interpolation, and the size of the low-frequency feature map is the same as the size of the first image.
[0048] Low-frequency information can reflect the general outline and background of the image, thus complementing the subsequent high-frequency enhancement methods.
[0049] Step S4: Enhance the high-frequency feature map. In this embodiment, the enhanced high-frequency feature map is: F EP =Concat(F H *Conv4,(F H *Conv2*CBAM*Conv3))
[0050] Wherein: F H For the high-frequency feature map before enhancement, F EP For the enhanced high-frequency feature map, Conv2 is the second convolution kernel, Conv3 is the third convolution kernel, Conv4 is the fourth convolution kernel, CBAM is the attention module operator, and Concat is the concatenation operation.
[0051] Compared to conventional convolution and pooling enhancement methods, this approach uses a combination of three convolutions and one attention layer to enhance high-frequency features, highlighting their details and clarity while completely ignoring contours and background. The enhanced high-frequency feature map and the upsampled low-frequency feature map are then concatenated and re-convolved to compensate for the lack of contours in the high-frequency feature map, thus enabling the identification of foreign objects suspended in various locations within the glass bottle.
[0052] In this embodiment, the size of the second, third, and fourth convolution kernels is 3*3.
[0053] Step S5: After fusing the enhanced high-frequency feature map and the upsampled low-frequency feature map, convolution is performed to obtain the output image. In this embodiment, this includes:
[0054] Step S5-1: The enhanced high-frequency feature map and the upsampled low-frequency feature map are stitched together along the channel dimension to obtain the first intermediate image;
[0055] Step S5-2: Perform a convolution operation on the first intermediate image using the fifth convolution kernel to obtain the output image. In this embodiment, the size of the fifth convolution kernel is also 3*3.
[0056] Furthermore, in some embodiments, the output channel can be adjusted using a 1x1 convolution.
[0057] Step S6: Input the output image into the target detection model to obtain the foreign object detection result.
[0058] All convolutional kernels and operators in steps S2 to S5 above are obtained through deep learning training. The training process is based on adjustments using a first loss function, which is:
[0059] Where: L is the first loss function. T(H) represents the normalized pixel value of the pixel in the i-th row and j-th column of the output image obtained in step S5. i,j ) represents the normalized pixel value of the pixel in the i-th row and j-th column of the output image in the training samples, and γ is the dynamic weight amplitude control parameter, which is set to be greater than 1 in this embodiment. It is an L1 norm.
[0060] In this embodiment, the normalized pixel value of any pixel is:
[0061] in: Let be the value of the k-th channel in the normalized pixel value of the pixel in the i-th row and j-th column. Let be the value of the k-th channel among the pixel values of the i-th row and j-th column before normalization. Let be the minimum value of all channels of the pixel in the i-th row and j-th column before normalization. It represents the maximum value of all channels of the pixel in the i-th row and j-th column before normalization.
[0062] In summary, by using X-ray sources with different voltages at the same viewing angle, all single-channel X-ray images obtained have the same viewing angle, making registration easier and avoiding new errors introduced by scaling and fusion. Furthermore, the use of different voltages can effectively cover the imaging needs of regions with different densities. The output images obtained by sequentially separating high and low frequencies, enhancing high frequencies, and then stitching and convolving them can clearly display foreign objects suspended in glass bottles. While improving detection accuracy and reliability, the foreign object detection process can be placed close to the filling process, eliminating the need for post-filling settling and greatly improving the overall production line efficiency.
[0063] Example 2
[0064] This embodiment is generally similar to Embodiment 1, with only a few differences. Specifically, the difference between this embodiment and Embodiment 1 is that in step S4 of this embodiment, the high-frequency feature map is enhanced twice, and the two enhancement methods are the same.
[0065] In addition, to further verify the effectiveness of this application, the following comparative examples are provided:
[0066] Comparative Example 1
[0067] This comparative example is designed with reference to Example 1. The difference between this comparative example and Example 1 is that in step S4, only convolution and pooling operations are used, and no splicing is performed. That is, the size of the enhanced high-frequency feature map is the same as that of the enhancement map. The rest is the same as Example 1.
[0068] Comparative Example 2
[0069] This comparative example is designed with reference to Chinese patent CN117351314A and combined with Chinese patent CN112508930A. A total of three different perspectives are arranged to synchronously collect data from the glass bottle under test. The light source used is an X-ray source. The arrangement direction of the three perspectives is the same as that of Chinese patent CN117351314A, and different voltages are used. The rest of the arrangement is based on CN117351314A.
[0070] By testing data from a glass-filled beverage production line, it was found that the time interval between the filling process and the foreign object detection process was approximately 6 seconds, leaving insufficient time for settling. A total of 1000 samples were obtained, with each sample consisting of three single-channel X-ray images. 700 samples were used as the training set, and the remaining 300 samples were used as the test set. The foreign object detection rate of Example 2 exceeded 97%, and that of Example 1 exceeded 91%. The detection rate of Comparative Example 2 was less than 10%, and that of Comparative Example 1 was approximately 60%.
[0071] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for detecting foreign objects inside a glass bottle, characterized in that, include: Step S1: For the glass bottle to be inspected, acquire multiple X-ray images of different voltages from the same viewing angle; Step S2: Align all the acquired X-ray images and synthesize them into a multi-channel first image; Step S3: After convolution and upsampling the first image to obtain the low-frequency feature map, subtract the low-frequency feature map from the first image to obtain the high-frequency feature map; Step S4: Enhance the high-frequency feature map; Step S5: After fusing the enhanced high-frequency feature map and the upsampled low-frequency feature map, perform convolution to obtain the output image; Step S6: Input the output image into the target detection model to obtain the foreign object detection result; In the enhancement process of step S4, the enhanced high-frequency feature map is as follows: F EP =Concat(F H *Conv4,(F H *Conv2*CBAM*Conv3)) Wherein: F H For the high-frequency feature map before enhancement, F EP For the enhanced high-frequency feature map, Conv2 is the second convolution kernel, Conv3 is the third convolution kernel, Conv4 is the fourth convolution kernel, CBAM is the attention module operator, and Concat is the concatenation operation. All convolutional kernels and operators in steps S2 to S5 are obtained through deep learning training. The training process is adjusted based on a first loss function, which is: Where: L is the first loss function. T(H) represents the normalized pixel value of the pixel in the i-th row and j-th column of the output image obtained in step S5. i,j ) represents the normalized pixel value of the pixel in the i-th row and j-th column of the output image in the training sample, and γ is the dynamic weight amplitude control parameter.
2. The method for detecting foreign objects inside a glass bottle according to claim 1, characterized in that, In step S1, three X-ray images with different voltages are acquired from the same viewpoint, and the first image is a three-channel image.
3. The method for detecting foreign objects inside a glass bottle according to claim 1, characterized in that, The first convolution kernel in step S3 has a size of 3*3 and a stride of 2.
4. The method for detecting foreign objects inside a glass bottle according to claim 1, characterized in that, The upsampling process in step S3 uses bilinear interpolation, and the size of the low-frequency feature map is the same as the size of the first image.
5. The method for detecting foreign objects inside a glass bottle according to claim 1, characterized in that, Step S5 includes: Step S5-1: The enhanced high-frequency feature map and the upsampled low-frequency feature map are stitched together along the channel dimension to obtain the first intermediate image; Step S5-2: Perform a convolution operation on the first intermediate image using the fifth convolution kernel to obtain the output image.
6. The method for detecting foreign objects inside a glass bottle according to claim 1, characterized in that, The normalized pixel value for any pixel is: in: Let be the value of the k-th channel in the normalized pixel value of the pixel in the i-th row and j-th column. Let be the value of the k-th channel among the pixel values of the i-th row and j-th column before normalization. Let be the minimum value of all channels of the pixel in the i-th row and j-th column before normalization. It represents the maximum value of all channels of the pixel in the i-th row and j-th column before normalization.
7. A foreign object detection device for a glass bottle, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-6.
8. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the method as described in any one of claims 1-6.