Cigarette packet surface defect detection method and system and storage medium

By combining multimodal light field control and depth estimation models with self-supervised learning for defect classification, the problems of single viewing angle and low precision in highly reflective cigarette package surface inspection are solved, achieving efficient and accurate defect detection.

CN120703116APending Publication Date: 2025-09-26CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510974697.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing technology for detecting surface defects on highly reflective cigarette packages has problems such as a single viewing angle and low detection accuracy. It is particularly difficult to maintain detection stability and accuracy in complex or dynamic environments.

Method used

A multimodal light field control model, distributed camera array and depth estimation model are adopted, combined with a self-supervised learning defect classification model, to perform defect detection through multi-angle images and multi-polarization angle multispectral images, and build a joint detection model.

Benefits of technology

The detection rate of tiny defects has been significantly improved, and the detection accuracy and robustness have been enhanced, ensuring efficient and accurate identification of defects on the surface of highly reflective cigarette packages in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120703116A_ABST
    Figure CN120703116A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial visual inspection, in particular to a cigarette packet surface defect detection method and system and a storage medium, and the method comprises the following steps: optimizing a light field by adopting a multi-mode light field regulation and control model; a distributed camera array is arranged, and a multi-angle image, a multi-mode image and a multi-polarization-angle multi-spectral image of the surface of a to-be-detected cigarette packet are obtained; according to the multi-angle image and the multi-modal image, obtaining a three-dimensional reconstruction depth map by adopting a depth estimation model; based on the multi-polarization-angle multispectral image, defect detection is carried out, and a defect edge is extracted; constructing a defect classification model based on self-supervised learning; carrying out joint training on the multi-modal light field regulation and control model, the depth estimation model and the defect detection and defect classification model to obtain a joint detection model; and performing defect detection on a to-be-detected cigarette packet by adopting the trained joint detection model. According to the embodiment of the invention, by fusing the dynamic adjustment light field, the deep learning algorithm and the multi-angle imaging, the defect detection performance in the high-reflection environment is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial visual inspection, and in particular to a method, system and storage medium for detecting surface defects of cigarette packages. Background Art

[0002] In cigarette production, defect detection on highly reflective packaging surfaces (such as laser-coated materials) presents significant technical challenges. Traditional optical inspection methods are susceptible to interference from specular reflections, resulting in image degradation and, in turn, impacting detection accuracy and reliability. While existing multispectral imaging and polarization imaging technologies can suppress reflections to a certain extent and provide clearer surface images, they are significantly limited in their dynamic adaptability and ability to adjust inspection strategies in real time.

[0003] Current technologies for defect detection on highly reflective surfaces typically rely on static light field manipulation and single-view image acquisition methods. While these methods can suppress reflections and capture surface defects, they still suffer from several technical shortcomings. First, the single-view light field manipulation. Existing technologies generally lack dynamic adaptive adjustment mechanisms for light sources and sensor angles, making it impossible to optimize detection strategies in real time based on lighting variations, surface material, or reflectivity in the production environment. This makes it difficult to maintain stable detection results in complex or dynamic environments. Especially on highly reflective surfaces or in highly variable production environments, lighting conditions and surface characteristics can affect the direction and intensity of reflected light. Without the ability to adjust the light source and sensor angle in real time, detection accuracy and stability will be significantly compromised. Second, image processing methods are insufficient. Existing multispectral and polarization imaging technologies mostly rely on a single viewpoint and static image processing, ignoring the differences in surface reflectivity at different observation angles. Without effectively combining multiple viewpoints or temporal information for image enhancement and defect extraction, these static processing methods fail to fully utilize the diversity of reflections and surface details, thus compromising defect detection accuracy. Especially for highly reflective surfaces, static images struggle to capture the complex optical properties of the surface, leading to overlooking or misidentifying subtle defects. Furthermore, defect recognition capabilities are limited. Most existing defect recognition models exhibit low detection accuracy when dealing with complex defects (such as tiny cracks and internal bubbles). In particular, they lack comprehensive multi-level, multi-scale analysis capabilities for identifying subtle defects. This limits the application of traditional detection methods in high-precision, high-efficiency industrial production environments. For complex defects such as tiny cracks or bubbles, single-view and static image processing methods often fail to capture sufficient detail, resulting in the inability to accurately identify these subtle defects, which in turn affects product quality and production efficiency.

[0004] The shortcomings of existing technologies limit their widespread practical application. More intelligent, multi-angle, and high-precision inspection methods are urgently needed to overcome these bottlenecks and improve the comprehensiveness and accuracy of defect detection. Therefore, for highly reflective cigarette packages (such as those made of laser-coated materials), it is imperative to introduce more intelligent dynamic adaptation mechanisms, more diverse image processing technologies, and more efficient defect recognition models to achieve high-precision inspection. Summary of the Invention

[0005] One of the purposes of the present invention is to provide a method, system and storage medium for detecting surface defects of cigarette packages, so as to solve the technical problems of the commonly used high-reflective surface defect detection methods in the prior art, such as the single viewing angle and low detection accuracy.

[0006] To achieve the above objectives, an embodiment of the present invention provides a method for detecting surface defects of cigarette packages, comprising: Adopting multimodal light field control model to optimize the light field; Setting up a distributed camera array and acquiring multi-angle images, multi-modal images, and multi-polarization angle multispectral images of the surface of the cigarette package to be inspected; Acquire a three-dimensional reconstructed depth map using a depth estimation model based on the multi-angle images and the multimodal images; Performing defect detection based on the multi-polarization angle multispectral image to extract defect edges; Build a defect classification model based on self-supervised learning; Jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model; The trained joint detection model is used to perform defect detection on the cigarette packs to be inspected.

[0007] Optionally, acquiring a three-dimensional reconstructed depth map using a depth estimation model according to the multi-angle image and the multimodal image includes: According to formula (1), a depth map is obtained from the multi-angle image. , (1) in, is a multi-angle depth map, is the output of the deep learning model, are the parameters of the model, is the multi-angle image; According to formula (2), the objective function is optimized. , (2) in, is the objective function, is the actual depth value.

[0008] Optionally, acquiring a three-dimensional reconstructed depth map using a depth estimation model according to the multi-angle image and the multimodal image includes: According to formula (3), the depth fusion map is obtained. , (3) in, To fuse the depth map, is a multimodal image, For joint deep learning networks; Calculate the loss function value according to formula (4): , (4) in, is the loss function value, is the actual depth value.

[0009] Optionally, acquiring a three-dimensional reconstructed depth map using a depth estimation model according to the multi-angle image and the multimodal image includes: According to formula (5), the 3D reconstruction depth is obtained. , (5) in, is the 3D reconstructed depth map, To fuse the depth map, is a multi-angle depth map, is the regularization coefficient, is a constraint item.

[0010] Optionally, performing defect detection based on the multi-polarization angle multispectral image to extract defect edges includes: According to formula (6) and formula (7), the fused image is obtained. , (6) , (7) in, To fuse the images, The polarization angle is The spectral band is images, is the weight function, For images In position The gradient amplitude of is a hyperparameter, is the normalization term; performing enhancement processing and noise reduction processing on the fused image; The edge detection algorithm is used to extract the defect edges from the processed fusion image.

[0011] Optionally, jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model includes: According to formula (8), the defect classification model is pre-trained through self-supervision. , (8) in, is the loss function, is the output of the defect classification model, is the image after occlusion, is the real image information of the occluded area, is the regularization hyperparameter, is a constraint item.

[0012] Optionally, jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model includes: The parameters of the light field are updated by gradient back propagation according to formula (9), , (9) in, is the optimized square control parameter vector, is the output of the defect classification model, is the defect classification loss function, is the light field control parameter vector, is the cross-modal fusion feature, is the regularization coefficient, is the regularization term.

[0013] Optionally, jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model includes: According to formula (10), self-supervised learning is guided. , (10) in, is the loss function, is the two-dimensional image reconstructed by the model, The 3D depth map reconstructed by the model, is a real two-dimensional image, is the real 3D depth map.

[0014] On the other hand, the present invention further provides a cigarette package surface defect detection system, the system comprising a processor, and the processor is configured to execute any of the above methods.

[0015] In another aspect, the present invention further provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed by a processor, any of the above methods is implemented.

[0016] Beneficial effects of the present invention: The embodiments of the present invention use multi-view image acquisition and adaptive light field control technology to jointly detect defects of different types and sizes, significantly improving the detection rate of small defects. By introducing advanced image processing algorithms and multi-level defect recognition technology, the robustness of the system in ultra-high-speed cigarette production equipment and complex environments is improved, and the real-time response capability and defect detection accuracy are enhanced. By optimizing light field control and image processing, the accuracy and reliability of detection are further improved, ensuring efficient and accurate defect detection under various environmental conditions. It solves the problem of defect detection in highly reflective environments, especially for the detection of defects on the surface of highly reflective cigarette packages made of laser materials, and the detection accuracy is improved by more than 30%.

[0017] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings: Figure 1 is a flow chart of a method for detecting surface defects of cigarette packages according to one embodiment of the present invention; Figure 2 A schematic diagram of the layout of a camera and a lighting lamp according to an embodiment of the present invention; Figure 3 A flowchart of a method for obtaining a three-dimensional reconstruction depth map according to one embodiment of the present invention; Figure 4 is a flowchart of a method for extracting defect edges according to one embodiment of the present invention; Figure 5 2 is a diagram of a joint training framework according to one embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.

[0020] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of laws and regulations. In the embodiments of this application, certain software, components, models, and other existing solutions in the industry may be mentioned. These should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use such solutions.

[0021] like Figure 1 The flowchart of the method for detecting surface defects of cigarette packs according to one embodiment of the present invention is shown. Figure 1 In the present invention, the detection method may include the following steps: In step S10, the light field is optimized using a multimodal light field control model; In step S11, a distributed camera array is set up to acquire multi-angle images, multi-modal images, and multi-polarization angle multispectral images of the surface of the cigarette pack to be inspected; In step S12, a depth estimation model is used to obtain a three-dimensional reconstructed depth map based on the multi-angle image and the multi-modal image; In step S13, defect detection is performed based on the multi-polarization angle multispectral image to extract the defect edge; In step S14, a defect classification model based on self-supervised learning is constructed; In step S15, the multimodal light field control model, the depth estimation model, the defect detection and defect classification model are jointly trained to obtain a joint detection model; In step S16, the trained joint detection model is used to perform defect detection on the cigarette pack to be inspected.

[0022] In this Figure 1 In the cigarette package surface defect detection method shown, step S10 is used to optimize the light field. Unlike the prior art that relies on a single control dimension, the embodiment of the present invention achieves fine adjustment of the light field through multimodal control, comprehensively optimizes the control effect of the light field, and thus improves the intensity and quality of the surface microstructure diffraction signal. In this embodiment, an adaptive feedback mechanism can be used to collect diffraction signal data in real time and compare it with the expected signal to automatically adjust the various control parameters of the light field (including phase, amplitude, polarization, etc.). Furthermore, the specific method for optimizing the light field in step S10 can be a variety of forms known to those skilled in the art. In one example of the present invention, step S10 may include: In step S20, the diffraction signal is collected and analyzed; In step S21, a feedback optimization algorithm is used to optimize light field parameters.

[0023] Step S20 is used to detect the intensity distribution of the diffracted light field, and the system dynamically adjusts the control parameters of the SLM according to the error between the detection result and the expected value. Step S21 is used to use intelligent optimization technologies such as particle swarm optimization (PSO) or deep learning algorithms to adjust the phase, amplitude and polarization state of the light field in real time. The feedback mechanism continuously optimizes the light field parameters to adapt to different surface microstructures and detection requirements. In response to the detection requirements of complex surface microstructures, in this example, an accurate diffraction signal calculation model is adopted to ensure high-precision signal enhancement effect. By introducing the Fresnel diffraction model to replace the traditional Fraunhofer diffraction approximation, the propagation and diffraction pattern of light waves on complex surface microstructures can be more accurately described. In order to avoid the existing technical barriers, a far-field diffraction signal intensity formula is proposed to more accurately describe the propagation of the diffraction signal. Specifically, in this example, step S20 can be, for example, using formula (11) to express the far-field diffraction signal intensity: , (11) in, is the observation point in the far field (angle is and , the diffraction signal intensity at distance z); is the electric field distribution on the microstructure surface S, which is the electric field modulated by the spatial light modulator (SLM); is the wave number of light, λ is the wavelength of light; are the position coordinates on the surface microstructure; is the unit vector pointing to the observation point (along the diffraction direction); is the distance from the surface microstructure point to the observation point; is the angular frequency of light; is the time variable; It is a tiny area element of the surface microstructure.

[0024] The specific method of phase control in step S10 may be, for example, to use formula (12) to perform phase control: , (12) in, To control the phase of the rear light field, is the initial phase of the light field, is a Zernike polynomial, is the adjustment coefficient. The phase of the incident light is precisely adjusted by a spatial light modulator (SLM), and the aberrations in the optical system are described using polynomials. The phase diagram is dynamically adjusted to optimize the light wavefront morphology and improve the quality of the diffraction signal. The method for amplitude control in step S10 can be to adjust the amplitude distribution of the SLM to optimize the intensity distribution of the light wave, thereby further improving the clarity and contrast of the diffraction signal. Combined with an adaptive control algorithm, the amplitude is automatically adjusted to meet the detection requirements of different surface microstructures. The method for polarization control in step S10 can be to achieve the polarization state of the light field by adjusting a polarizer or a liquid crystal light modulator (LCM). By controlling the polarization direction and degree of polarization, the interference of reflected light is reduced and the detection accuracy of surface microstructures is improved.

[0025] Step S10 uses a multimodal control method to comprehensively adjust phase, amplitude, and polarization state, providing greater flexibility and precision. Furthermore, the use of a precise propagation model avoids errors introduced by simplified models and better accounts for lightwave propagation characteristics. Through dynamic optimization and real-time feedback mechanisms, light field parameters (such as phase, amplitude, and polarization state) can be adjusted in real time, overcoming the limitations of traditional static optimization and improving the performance of diffraction signals. This overcomes the drawback of traditional technologies that rely solely on polarization control.

[0026] Step S11 is used to set up a distributed camera array and obtain multi-angle images, multi-modal images and multi-polarization angle multispectral images of the surface of the cigarette package to be inspected. Traditional three-dimensional reconstruction usually relies on a fixed-angle camera array or a single perspective for data acquisition. However, these methods often have difficulty in providing sufficient coverage in scenes with complex surfaces or large changes in lighting, resulting in a decrease in the accuracy of defect detection. In this embodiment, a distributed camera array is used. The distributed camera array adopts a distributed layout of multiple cameras, and uses multiple perspectives at different positions and angles to simultaneously capture images, further increasing the coverage and resolution of surface defects. Unlike traditional annular or fixed arrays, the layout of the distributed camera array is more flexible and adaptive, and the camera can be dynamically adjusted as needed to ensure the best imaging effect in different environments, such as Figure 2 As shown. The layout and angle of the camera array are dynamically optimized to adapt to different detection scenarios. Specifically, in this example, the camera arrangement model based on optimization can be expressed by formula (13): , (13) in, It is The spatial position of the camera, It is The direction of the camera, It is The blind spot area of ​​the camera, is the number of cameras. Assume that at a given moment, the position and pose of the target object are known, and there are multiple defect areas on the target surface. Using an optimization algorithm (such as a genetic algorithm or particle swarm optimization), the position and orientation of each camera can be dynamically determined based on the characteristics of the target surface. The goal is to maximize the coverage of the camera array and minimize blind spots. This optimization algorithm dynamically adjusts the camera array layout based on the defect detection task, ensuring coverage of each area and avoiding blind spots and overlapping areas, thereby improving detection accuracy.

[0027] After optimizing the light field and setting up the distributed camera array, light field-camera array collaborative calibration is performed. In this embodiment, geometric calibration enhancement and polarization-spectral alignment optimization can be performed. Specifically, in this example, the specific method of geometric calibration enhancement can be to adopt a multi-coordinate system conversion model based on Zhang Zhengyou's calibration method, and establish a conversion relationship between the light field coordinate system (OL) and the camera coordinate system (O-Ci, i=1..N) through a calibration plate. For example, the conversion relationship of formula (14) can be established: , (14) in, is the camera coordinate system The three-dimensional point coordinates under is the three-dimensional point coordinate in the light field coordinate system, Light Field to Camera The rotation matrix (3×3 orthogonal matrix), is the translation vector (3D vector) from the light field to camera i, which is solved by nonlinear optimization, and the error convergence threshold is set to the pixel level (<1px).

[0028] The specific method for optimizing polarization-spectral alignment can be to establish a mapping relationship between the polarization data (degree of linear polarization DoLP) controlled by the light field and the band response function (BRF) of the multispectral camera. Specifically, the mapping relationship of formula (15) can be established: , (15) in, is the band response function of polarization angle θ at wavelength λ, is the linear polarization degree corresponding to the polarization angle θ (0-1), is the transmittance of the polarizer for light of wavelength λ. By minimizing the characteristic differences between different modes (e.g., maximizing mutual information), the wavelength selection and polarizer angle configuration of the spectral camera are optimized.

[0029] Step S12 is used to obtain a 3D reconstructed depth map using a depth estimation model based on the multi-angle images and the multi-modal images. Traditional disparity calculation relies on geometric formulas. In this embodiment, depth information is directly extracted from the multi-view images using a deep learning algorithm, avoiding the complex geometric calculations in the prior art. The specific method for obtaining the 3D reconstructed depth map in step S12 can be various forms known to those skilled in the art. In one example of the present invention, step S12 may include the following: Figure 3 The steps shown in Figure 3 In the step S11, the following steps may be included: In step S30, a multi-angle depth map is obtained using a deep learning model based on the multi-angle images; In step S31, a fused depth map is obtained using a joint deep learning model based on the multimodal image; In step S32, based on the multi-angle depth map and the fused depth map, a deep learning optimization strategy is adopted to obtain the final 3D reconstructed depth map.

[0030] In this Figure 3 In the method shown, step S30 is used to obtain a multi-angle depth map. In this embodiment, the specific method for obtaining the multi-angle depth map in step S30 can be various forms known to those skilled in the art. In one example of the present invention, step S30 can include the following steps: In step S40, a depth map is obtained from the multi-angle image according to formula (1): , (1) in, is a multi-angle depth map, is the output of the deep learning model, are the parameters of the model, is the multi-angle image; In step S41, the objective function is optimized according to formula (2): , (2) in, is the objective function, is the actual depth value.

[0031] Step S40 is used to adopt a deep learning network From the image Step S41 is used to optimize the target function, the goal of which is to train the network so that the depth map output by the model is With the real depth value Closest.

[0032] Step S31 is used to obtain a fused depth map based on the multimodal images using a joint deep learning model. By combining image data from multiple modalities (multispectral images, infrared images, etc.), the accuracy of depth estimation can be further improved. Therefore, in this embodiment, data obtained by different sensors is fused to generate more accurate depth information. Furthermore, the specific method for obtaining the fused depth map in step S31 can be various forms known to those skilled in the art. In one example of the present invention, step S31 may include: In step S50, the depth fusion map is obtained according to formula (3): , (3) in, To fuse the depth map, is a multimodal image, It is a joint deep learning network that estimates the final depth value by learning the relationship between images of different modalities.

[0033] In step S51, the loss function value is calculated according to formula (4): , (4) in, is the loss function value, is the actual depth value.

[0034] After obtaining the depth maps from multiple perspectives, in order to more accurately reconstruct the three-dimensional structure of the target object, step S32 can be used to fuse the depth maps from different perspectives to obtain the final three-dimensional reconstructed depth map. Specifically, in this example, step S32 can be, for example, using formula (5) to obtain the final three-dimensional reconstructed depth map: , (5) in, is the 3D reconstructed depth map, To fuse the depth map, is a multi-angle depth map, is the regularization coefficient, is a constraint item.

[0035] After the light field-camera array collaborative calibration, the light field-array data is preprocessed, including the spatiotemporal alignment pipeline and feature cascade fusion. In this example, the spatiotemporal alignment pipeline can be specifically implemented by first performing spatial coordinate transformation: converting the diffraction signal intensity generated by the light field manipulation to Mapped to the pixel coordinates of each camera coordinate system through the geometric calibration matrix , generating light field-visual joint features Secondly, time synchronization resampling is performed: light field parameters (such as polarization angle θ and spectral band λ) are interpolated and aligned with the camera image based on the timestamp to ensure temporal consistency of the data during fusion. Then, a cross-modal parameter mapping table is generated to establish the correspondence between light field control parameters (such as polarization angle θ and spectral band λ) and camera acquisition channels, as shown in formula (16).

[0036] , (16) in, The polarization angle of the camera and spectral bands The corresponding acquisition channel signal is as follows: Indicates that the spatial light modulator is based on the polarization angle The generated polarization control signal, Indicates the wavelength of the spectral filter Transmittance or selectivity characteristics under Represents a channel joint map.

[0037] In this example, the specific method of feature cascade fusion can be to design a two-stream convolutional neural network (CNN) to add light field-visual feature joint extraction. Specifically, it can first extract frequency domain features through 2D convolution based on the input diffraction signal characteristics (phase spectrum, polarization degree). Secondly, according to the input multi-view image With Depth Map , extracting spatial features through 3D convolution Then, through channel splicing and attention mechanism to generate cross-modal features For example, formula (17) can be used to generate cross-modal features. : , (17) By introducing multi-scale feature fusion, convolution kernels with different expansion rates (3×3 and 5×5) are used to extract multi-scale diffraction features from the light field branch, which are then pyramidally fused with the multi-view image features of the visual branch (such as shallow edge features and deep semantic features).

[0038] Step S13 is used to perform defect detection based on the multi-polarization angle multispectral image to extract the defect edge. In this embodiment, the specific method for extracting the defect edge in step S13 can be various forms known to those skilled in the art. In one example of the present invention, step S13 may include the following: Figure 4 The steps shown in Figure 4 In the step S13, the following steps may be included: In step S60, a fused image is acquired based on the multi-polarization angle multispectral image; In step S61, the fused image is enhanced and denoised; In step S62, an edge detection algorithm is used to extract defect edges from the processed fused image.

[0039] In this Figure 4 In the method shown, step S60 introduces an adaptive weighting function to automatically adjust the weight based on image quality or defect information, avoiding manual setting in traditional methods. Specifically, in this example, step S60 can be, for example, using formula (6) and formula (7) to obtain a fused image: , (6) , (7) in, To fuse the images, The polarization angle is The spectral band is images, is the weight function, For images In position The gradient amplitude of is a hyperparameter, is the normalization term.

[0040] Step S61 is used to perform enhancement and noise reduction on the fused image. By performing enhancement and noise reduction operations on the fused image, tiny defect details are highlighted and noise is removed. In this embodiment, an adaptive histogram equalization method (AHE) is used for image enhancement, and contrast is enhanced by histogram equalization of local areas, especially in low-contrast areas. Specifically, in this example, step S61 can be, for example, using formulas (18) and (19) to perform image enhancement and noise reduction: , (18) , (19) in, For adaptive histogram equalization, the Clip function is used to ensure that the enhanced value is within a reasonable range. Represents the minimum value of the image in the local area, Indicates the maximum value of the image in a local area. represents the loss function of the denoising model, For the denoised image, the goal is to minimize the noise in the image.

[0041] In the enhanced and denoised image, the edge of the defect is extracted using an edge detection algorithm in step S62. The edge region is determined by calculating the gradient and gradient direction of the image, and a double threshold process is applied. Finally, an edge map is generated by connecting the edge points. Specifically, in this example, step S62 can be, for example, to obtain the defect edge using formula (20): , (20) in, Indicates the location Is there an edge at the location? This value is based on the gradient size and a predetermined threshold. and to decide by comparison. and The threshold is set to distinguish strong edges, weak edges and non-edge areas. Indicates that if the point is between the low and high thresholds, but is connected to other strong edges, it is considered to be part of an edge.

[0042] Step S14 is used to build a defect classification model based on self-supervised learning. Self-supervised learning can automatically find potential defect areas from a large number of unlabeled images without labeled data by generating self-labeled data. Its core is to design a suitable loss function so that the model can learn from the self-construction of the image. Assume that there are a large number of unlabeled images. ,in Representative An image is obtained and the defect features in the image are identified through self-supervised learning. To achieve this goal, in this embodiment, a self-supervised loss function is designed so that the model learns by reconstructing or predicting the information of the defect area. Specifically, in this example, step S14 can be, for example, using formula (8) to perform self-supervised pre-training on the defect classification model. , (8) in, is the loss function, is the output of the defect classification model, is the image after occlusion, is the real image information of the occluded area, is the regularization hyperparameter, is a constraint. , randomly select a part of the image area to block, and get the blocked image The model then needs to predict information about the occluded areas, specifically the defect areas. The goal of the loss function is to enable the model to learn the underlying patterns of image defects through self-supervision, thereby automatically identifying defect features without labeled data. As the model's image reconstruction becomes increasingly accurate, it is able to identify defects from the underlying image features. At this point, the labels generated by the model (i.e., the predicted defect areas) can be used to further train the supervised learning model.

[0043] Step S14 is also used to design a conditional self-supervisory model, taking the light field parameters as additional input dimensions, forcing the model to learn defect characteristics under specific light field conditions. Specifically, the conditional self-supervisory model can be expressed, for example, using formula (21): ,(twenty one) in, represents the missing reconstructed image predicted by the model, Indicates that the parameter is The neural network model, represents a partially occluded or corrupted input image, is the light field control parameter vector, including phase coefficient, polarization angle and amplitude modulation value.

[0044] Then, the light field consistency constraint is added to the loss function. Specifically, for example, formula (22) can be used as the loss function: + ,(twenty two) in: is the light field phase distribution predicted by the model, represents the real light field phase distribution, is the polarization state of the light field predicted by the model, Represents the true polarization state of the light field. The loss function is used to ensure that the light field characteristics of the reconstructed image are consistent with the original control parameters.

[0045] In this embodiment, when there are fewer defect samples, GAN can be used to generate more defect images to enhance the diversity of the training set, thereby improving the generalization ability of the model. Step S14 can also include using a generative adversarial network (GAN) to generate training data. Specifically, in this example, step S14 can be, for example, using formula (23) to train the generative adversarial network to obtain defect image training data. ,(twenty three) in, To counter the loss function, is the defect image generated by the generator, is the output of the discriminator, is the expected value of the log probability of the discriminator for the real image, is the expected value of the logarithmic probability of the discriminator for the generated image. During training, the generator attempts to maximize the discriminator's probability of misclassification, making the generated defect images increasingly realistic, while the discriminator attempts to correctly distinguish between generated images and real images. Defect images generated by GANs can enhance the training dataset, allowing the model to see more defect samples during training, thereby improving the model's generalization ability to unseen defects.

[0046] Step S15 is used to jointly train the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model. The joint training framework is shown in FIG. Figure 5 In this embodiment, the specific method of the joint training in step S15 can be various forms known to those skilled in the art. In one example of the present invention, step S15 can include: In step S70, light field control parameters are optimized by using cross-module parameter sharing; In step S71, data enhancement is performed using the three-dimensional reconstructed depth map; In step S72, the defect detection and defect classification models are jointly trained.

[0047] Step S70 is used to optimize the light field control parameters by using cross-module parameter sharing. Specifically, in this example, the phase adjustment coefficient , polarization angle And the classification loss function Linkage, update the light field control model parameters through gradient back propagation, forming a closed loop of "defect classification feedback → light field parameter adaptive adjustment". Specifically, it can be, for example, to update the light field parameters through gradient back propagation using formula (9), , (9) in, is the optimized square control parameter vector, is the output of the defect classification model, is the defect classification loss function, is the light field control parameter vector, is the cross-modal fusion feature, is the regularization coefficient, is the regularization term.

[0048] Step S71 is used to perform data enhancement using the 3D reconstructed depth map. Specifically, in this example, in the occlusion reconstruction task of self-supervised learning, the combined 3D reconstructed depth map Perform three-dimensional space occlusion: three-dimensional point cloud Randomly generate spherical occlusion areas; require the model to simultaneously reconstruct the missing areas of the two-dimensional image and the three-dimensional depth values. Specifically, formula (10) can be used as the loss function to guide self-supervised learning: , (10) in, is the loss function, is the two-dimensional image reconstructed by the model, The 3D depth map reconstructed by the model, is a real two-dimensional image, The real 3D depth map is obtained by forcing the model to learn the spatial correlation of defects, thereby improving the detection capability of 3D defects (such as bubbles and pits). Step S72 is used to jointly train the defect detection and defect classification models.

[0049] Step S16 is used to use the trained joint detection model to perform defect detection on the cigarette pack to be inspected.

[0050] On the other hand, the present invention further provides a cigarette package surface defect detection system, the system comprising a processor, and the processor is configured to execute any of the above methods.

[0051] In another aspect, the present invention further provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed by a processor, any of the above methods is implemented.

[0052] Beneficial effects of the present invention: The embodiments of the present invention adopt a dynamic light field control mechanism. By dynamically adjusting the wavefront and polarization state of the light source according to real-time changes in the production line environment (such as light intensity, cigarette package material, equipment position, etc.), it greatly reduces the impact of the external environment on the detection results and ensures high-precision defect detection under different working conditions.

[0053] The embodiments of the present invention use reflection suppression and image enhancement technology, polarization image fusion and multispectral feature extraction methods to effectively suppress specular reflection and highlight diffuse reflection features, so that the detection system can more clearly capture tiny surface flaws and internal defects.

[0054] The embodiments of the present invention adopt multi-view imaging and high-frame-rate synchronous acquisition. By precisely arranging multiple high-frame-rate industrial cameras, it covers multiple angles of the object and synchronously acquires images in real time, eliminating the blind spots that may be caused by a single viewpoint, thereby further improving the comprehensiveness and accuracy of detection.

[0055] The embodiments of the present invention introduce a multi-level defect analysis framework, which can effectively handle defects of different scales and types, especially has outstanding advantages in the detection of tiny defects, and is suitable for the accurate identification and differentiation of complex surface defects.

[0056] The implementation method of the present invention uses GAN to generate real defect samples in a highly reflective environment, making up for the problem of insufficient defect data. At the same time, it combines the attention mechanism classifier to accurately identify subtle defects in the laser area, automatically ignore reflective artifacts, and improve the accuracy and robustness of defect recognition.

[0057] The implementation of the present invention adopts self-supervised learning technology, which reduces the dependence on manually labeled data. It not only reduces the cost of labeling and training, but also improves the system's adaptability and self-optimization capabilities, ensuring that the system can continuously improve its performance and adapt to different production conditions.

[0058] By integrating advanced technologies such as dynamic light field adjustment, deep learning algorithms, and multi-angle imaging, this invention optimizes defect detection performance in highly reflective environments, significantly improving detection accuracy and real-time performance, and enhancing the system's robustness and adaptability in complex environments. This system can effectively reduce manual intervention, improve production efficiency, and reduce maintenance costs, possessing broad application prospects and economic value.

[0059] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0061] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0063] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0064] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0065] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0066] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0067] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for detecting surface defects of cigarette packages, characterized in that: The detection method comprises: Adopting multimodal light field control model to optimize the light field; Setting up a distributed camera array and acquiring multi-angle images, multi-modal images, and multi-polarization angle multispectral images of the surface of the cigarette package to be inspected; Acquire a three-dimensional reconstructed depth map using a depth estimation model based on the multi-angle images and the multimodal images; Performing defect detection based on the multi-polarization angle multispectral image to extract defect edges; Build a defect classification model based on self-supervised learning; Jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model; The trained joint detection model is used to perform defect detection on the cigarette packs to be inspected.

2. The detection method according to claim 1, wherein Acquiring a three-dimensional reconstructed depth map using a depth estimation model according to the multi-angle image and the multi-modal image includes: According to formula (1), a depth map is obtained from the multi-angle image. ,(1) in, is a multi-angle depth map, is the output of the deep learning model, are the parameters of the model, is the multi-angle image; According to formula (2), the objective function is optimized. ,(2) in, is the objective function, is the actual depth value.

3. The detection method according to claim 1, wherein Acquiring a three-dimensional reconstructed depth map using a depth estimation model according to the multi-angle image and the multi-modal image includes: According to formula (3), the depth fusion map is obtained. ,(3) in, To fuse the depth map, is a multimodal image, For joint deep learning networks; Calculate the loss function value according to formula (4): ,(4) in, is the loss function value, is the actual depth value.

4. The detection method according to claim 1, wherein Acquiring a three-dimensional reconstructed depth map using a depth estimation model according to the multi-angle image and the multi-modal image includes: According to formula (5), the 3D reconstruction depth is obtained. ,(5) in, is the 3D reconstructed depth map, To fuse the depth map, is a multi-angle depth map, is the regularization coefficient, is a constraint item.

5. The detection method according to claim 1, wherein Performing defect detection based on the multi-polarization angle multispectral image to extract defect edges includes: According to formula (6) and formula (7), the fused image is obtained. ,(6) ,(7) in, To fuse the images, The polarization angle is The spectral band is images, is the weight function, For images In position The gradient amplitude of is a hyperparameter, is the normalization term; performing enhancement processing and noise reduction processing on the fused image; The edge detection algorithm is used to extract the defect edges from the processed fusion image.

6. The detection method according to claim 1, characterized in that Jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model includes: According to formula (8), the defect classification model is pre-trained through self-supervision. ,(8) in, is the loss function, is the output of the defect classification model, is the image after occlusion, is the real image information of the occluded area, is the regularization hyperparameter, is a constraint item.

7. The detection method according to claim 1, characterized in that Jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model includes: The parameters of the light field are updated by gradient back propagation according to formula (9), ,(9) in, is the optimized square control parameter vector, is the output of the defect classification model, is the defect classification loss function, is the light field control parameter vector, is the cross-modal fusion feature, is the regularization coefficient, is the regularization term.

8. The detection method according to claim 1, wherein Jointly training the multimodal light field control model, the depth estimation model, the defect detection and defect classification model to obtain a joint detection model includes: According to formula (10), self-supervised learning is guided. ,(10) in, is the loss function, is the two-dimensional image reconstructed by the model, The 3D depth map reconstructed by the model, is a real two-dimensional image, is the real 3D depth map.

9. A cigarette package surface defect detection system, characterized in that: The system comprises a processor configured to execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Acoustic module defect classification method and system based on multi-feature fusion

    CN121280813A

  • A Defect Classification Method and System for Acoustic Modules Based on Multi-Feature Fusion

    CN121280813B