A Ceramic Packaging Detection Method Based on Autoencoder and Improved ResNet18 Model

Through the autoencoder combined with the improved ResNet18 model, the detection accuracy instability caused by lighting changes and equipment jitter in ceramic packaging detection is solved, and efficient and accurate ceramic missed detection is achieved to adapt to complex industrial environments.

CN119831991BActive Publication Date: 2025-07-22HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510305282.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-22
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The prior art faces challenges such as lighting changes, equipment jitters and color interference in ceramic packaging inspection, resulting in unstable detection accuracy, especially in complex industrial environments, and is difficult to accurately judge the packaging status, and it is highly dependent on a large amount of labeled data.

Method used

The autoencoder is used to combine the CBAM module with the improved ResNet18 model. By constructing the ceramic packaging box image database, the autoencoder is used for image matching and segmentation, and feature extraction is performed with the improved residual unit and attention mechanism to realize ceramic missed installation detection.

Benefits of technology

It improves the reliability and efficiency of detection, reduces dependence on labeled data, enhances the recognition accuracy of the model in lighting changes and occlusion scenarios, and ensures the stable and efficient operation of the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831991B_ABST
    Figure CN119831991B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of target detection, and discloses a ceramic packaging detection method based on an autoencoder and an improved ResNet18 model, including the following steps: a database construction step of constructing a ceramic packaging box image database; a model construction step of combining the improved ResNet18 and CBAM to construct a product packaging box status detection model; an image acquisition step of acquiring an image of the ceramic packaging box to be detected as a detection image, and using the autoencoder model to match the detection image with the ceramic packaging box image database to obtain position information; an image segmentation step of using the position information to perform regional segmentation on the detection image to extract the sub-images to be detected; a packaging detection step of inputting the sub-images into the product packaging box status detection model to detect whether the ceramics are missing from the package. The present invention demonstrates excellent adaptability and robustness in a complex industrial environment, laying a solid foundation for the accuracy and reliability of intelligent packaging detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a ceramic packaging detection method based on an autoencoder and an improved ResNet18 model. Background Art

[0002] In the ceramic packaging detection process, the system mainly faces challenges such as light changes, equipment jitter, and color interference. Due to the unstable light conditions in the factory environment, the lighting may change over time or position. Coupled with the slight jitter of the equipment and the deviation of the shooting angle, the images of the same packaging box taken at different times show obvious differences. These changes easily lead to inaccurate matching between the sample images and the database templates, affecting the detection accuracy. In addition, when the color of the ceramic is similar to the background color of the packaging box, the boundary between the object and the background becomes blurred, and it is difficult for the model to effectively distinguish between the two. This visual confusion can lead to missed detections or false detections. Especially in the scenario of high-frequency detection on the production line, frequent errors may lead to a decrease in production efficiency. The light fluctuations and the reflection on the ceramic surface will also form highlight areas or shadows in the image, further interfering with the system's capture of the object contour. In such a complex environment, the model needs to have strong feature extraction and generalization capabilities to ensure that it can accurately judge the packaging status under the conditions of close colors, unstable lighting, or occlusion interference, and issue an alarm in time when there is a missing installation to ensure the stable and efficient operation of the production line.

[0003] The application of existing methods in ceramic packaging detection still has many deficiencies. Manual detection is inefficient, and long-term operation is prone to fatigue, which increases the risk of missed detections and false detections. At the same time, the judgment criteria among different operators are inconsistent, which also leads to the instability of the detection results. Although deep learning technology has shown certain potential, it highly depends on a large amount of labeled data, which makes the data collection and labeling costs remain high in actual production. In addition, the model performs limitedly in dealing with complex scenarios such as light changes, ceramic reflection, and object occlusion, and it is difficult to maintain consistent detection accuracy. The problem of class imbalance further exacerbates the difficulty of missing installation detection because the model tends to predict common classes and ignores abnormal situations. These deficiencies reflect the lack of generalization ability and adaptability of the current methods, and there is an urgent need for further optimization to meet the detection requirements in complex industrial environments. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art.

[0005] The technical solution adopted by the present invention to solve its technical problems is: to provide a ceramic packaging detection method based on an autoencoder and an improved ResNet18 model, including the following steps:

[0006] Database construction step, constructing a ceramic packaging box image database;

[0007] Model construction step: construct a product packaging box status detection model by combining the improved ResNet18 and CBAM;

[0008] Image acquisition step: acquire an image of the ceramic packaging box to be detected as the detection image, and use the autoencoder model to match the detection image with the ceramic packaging box image database to obtain the position information;

[0009] Image segmentation step: use the position information to perform regional segmentation on the detection image and extract the sub-images to be detected;

[0010] Packaging detection step: input the sub-images into the product packaging box status detection model to detect whether there is a shortage of ceramics; if there is no shortage, the detection ends; if there is a shortage, replenish the ceramics and re-enter the image acquisition step;

[0011] The product packaging box status detection model includes a CBR module, a max pooling layer, a first improved residual unit, a first CBAM module, a second improved residual unit, a second CBAM module, a Conv AvgPool module, and a fully connected layer connected in sequence; the CBR module consists of convolution, batch normalization, and a ReLU activation function, and is used to extract features from the input image; the max pooling layer is used to downsample the features extracted by the CBR module; the first improved residual unit includes an improved residual block, which is used to further extract features from the features output by the max pooling layer; the first CBAM module includes channel attention and spatial attention mechanisms, which are used to dynamically adjust the weights of important channels and regions in the feature map output by the first improved residual unit; the second improved residual unit includes several improved residual blocks, which are used to further extract features from the feature map output by the first CBAM module multiple times; the second CBAM module includes channel attention and spatial attention mechanisms, which are used to dynamically adjust the weights of important channels and regions in the feature map output by the second improved residual unit; the Conv AvgPool module further extracts features from the feature map output by the second CBAM module through convolution operations and downsamples the extracted features through average pooling; the fully connected layer converts the features output by the Conv AvgPool module into a probability distribution, thereby outputting the classification result of whether there is a shortage of ceramics;

[0012] The improved residual block includes a first residual layer and a second residual layer connected in sequence. The first residual layer adopts a dual-path interaction mechanism. One path quickly extracts shallow texture features through a single ALC, and the other path deeply mines semantic information through two cascaded ALCs. Finally, the extraction results of the two paths are added together and then passed through the ReLU activation function; the second residual layer adopts a single-path structure. The input features are directly processed by two cascaded ALCs and then added to the input features through a skip connection, and then passed through the ReLU activation function;

[0013] The ALC includes a variable batch normalization, a LeakyReLU activation function, a first convolutional layer, a variable batch normalization, a LeakyReLU activation function, and a second convolutional layer connected in sequence. The input of the ALC is added to the output of the second convolutional layer through a skip connection to obtain the final output of the ALC;

[0014] The variable batch normalization is expressed as:

[0015] ; where represents the original input, represents the result of the variable batch normalization; and are the mean and variance of the current batch, and are dynamically generated parameters; is a very small positive number used to prevent division by zero.

[0016] Preferably, the construction of the ceramic packaging box image database includes the following steps:

[0017] Collect images of ceramic packaging boxes without ceramics;

[0018] Manually annotate the collected images, mark the empty spaces in the packaging box where ceramics need to be loaded, and perform auxiliary box annotation.

[0019] Preferably, the CBAM module performs the following steps on the input feature map:

[0020] For the input feature map , calculate the channel attention weight and the spatial attention weight , respectively, which are expressed as:

[0021] ;

[0022] ;

[0023] where is the Sigmoid function, represents the concatenation operation on the channel dimension; MLP represents the multi-layer perceptron operation, AvgPool represents the average pooling operation; MaxPool represents the maximum pooling operation; Conv 7×7 represents the convolution operation using a 7×7 convolution kernel;

[0024] Multiply the channel attention weight by the original feature map to obtain the channel-weighted feature map , which is expressed as:

[0025] ;

[0026] Among them, represents element-wise multiplication;

[0027] Multiply the spatial attention weight with the channel-weighted feature map to obtain the final output feature map , which is expressed as:

[0028] .

[0029] Preferably, the loss function of the product packaging box status detection model is expressed as:

[0030] ;

[0031] ;

[0032] ;

[0033] Among them, represents the loss function of the product packaging box status detection model; represents the cross-entropy loss. The total number of samples is N, and the total number of categories is C. represents whether sample i belongs to category c. If so , otherwise ; is the predicted probability that sample i belongs to category c; is the regularization term, is the regularization coefficient, which is used to control the strength of regularization. represents the total number of all trainable parameters in the model. represents the parameter index, is the model parameter.

[0034] Preferably, using the autoencoder to match the detection image with the ceramic packaging box image database includes the following steps:

[0035] Use the encoder to compress the input detection image into a feature vector and reconstruct a high-quality image;

[0036] Compare the feature vector generated by the encoder with the images in the ceramic packaging box image database, identify the most matching image, and extract the position information therein;

[0037] The decoder uses the feature vector output by the encoder to reconstruct an output image with the same size as the input detection image.

[0038] Preferably, the encoder is used to compress the input detection image into a feature vector. Specifically, a 2D convolutional layer is adopted to extract features from the detection image, and the SENet channel attention mechanism is introduced during feature extraction to calculate the channel weights, which is expressed as:

[0039] ;

[0040] where z represents the original feature map, and are the learned parameters, represents the LeakyReLU activation function, represents the sigmoid activation function;

[0041] The weight s is multiplied by the original feature map to complete feature recalibration.

[0042] Preferably, the decoder uses the feature vector output by the encoder to reconstruct an output image with the same size as the input detection image. Specifically: the feature vector is expanded into a larger feature map through a fully connected layer, and then a transposed convolution operation is used to convert the feature map into the same size as the input image.

[0043] Preferably, comparing the feature vector generated by the encoder with the images in the ceramic packaging box image database is specifically achieved by calculating the cosine similarity between the feature vector generated by the encoder and the image feature vectors in the ceramic packaging box image database. The calculation formula of the cosine similarity is:

[0044] ;

[0045] where A and B are two feature vectors.

[0046] Preferably, the loss function of the autoencoder is:

[0047] ;

[0048] where n represents the number of pixel points, represents the actual value of the i-th pixel point, represents the predicted value of the i-th pixel point output by the decoder.

[0049] The present invention has the following beneficial effects:

[0050] (1) The present invention converts the packaging box image into a feature vector through an autoencoder and matches the samples in the database, effectively reducing the interference of factors such as illumination changes and shooting angle differences on the detection. This process ensures the accuracy of the auxiliary box annotation and the precise matching of positions, enabling subsequent ceramic missing detection to be carried out in a more complex environment, thus significantly improving the reliability of the detection.

[0051] (2) In the aspect of feature extraction, the present invention introduces an improved ResNet18 and CBAM module, enhancing the overall performance of the model. ResNet18 can capture multi-level visual features through its deep residual structure, while the CBAM module strengthens the focusing ability on key features through channel attention mechanism and spatial attention mechanism. This combination significantly improves the recognition accuracy of the model in complex scenarios such as lighting changes, object overlaps, and visual similarities, ensuring the robustness of the system in practical applications.

[0052] (3) The present invention reduces the dependence on a large amount of labeled data. Through the lightweight autoencoder design and feature optimization, it not only improves the detection efficiency but also can respond to the needs of the production line in real time, enabling the ceramic packaging detection to more quickly adapt to different production conditions, thereby improving the overall production efficiency and accuracy.

[0053] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Brief Description of the Drawings

[0054] Figure 1 It is the method step diagram of the embodiment of the present invention;

[0055] Figure 2 It is the process schematic diagram of the embodiment of the present invention;

[0056] Figure 3 It is the schematic diagram of the ceramic packaging box image of the embodiment of the present invention; among them, (a) is the image of an empty ceramic packaging box, and (b) and (c) are the images of an empty ceramic packaging box marked with auxiliary frames;

[0057] Figure 4 It is the network structure diagram of the autoencoder model of the embodiment of the present invention;

[0058] Figure 5 It is the network structure diagram of the product packaging box status detection model of the embodiment of the present invention;

[0059] Figure 6 It is the ALC network structure diagram of the embodiment of the present invention;

[0060] Figure 7 It is the CBAM attention mechanism diagram of the embodiment of the present invention;

[0061] Figure 8 It is the packaging detection schematic diagram of the embodiment of the present invention;

[0062] Figure 9 It is the test effect diagram of the embodiment of the present invention. Specific Embodiments

[0063] See Figure 1As shown in the figure, it is the method step diagram of the embodiment of the present invention, including the following steps:

[0064] S101, Database construction step, constructing a ceramic packaging box image database;

[0065] S102, Model construction step, constructing a product packaging box status detection model by combining the improved ResNet18 and CBAM;

[0066] S103, Image acquisition step, acquiring an image of the ceramic packaging box to be detected as a detection image, and using an autoencoder model to match the detection image with the ceramic packaging box image database to obtain position information;

[0067] S104, Image segmentation step, using the position information to perform regional segmentation on the detection image and extracting the sub-images to be detected;

[0068] S105, Packaging detection step, inputting the sub-images into the product packaging box status detection model to detect whether there is any missing ceramic; if there is no missing ceramic, the detection ends, and if there is a missing ceramic, the ceramic is replenished and the image acquisition step is re-entered.

[0069] See Figure 2 As shown in the figure, the entire working process of ceramic packaging detection can be divided into two key stages.

[0070] Specifically, Stage 1 includes database establishment, auxiliary box annotation, and position matching.

[0071] Database establishment. To construct a database for the packaging detection system, the embodiments of the present invention collect images of various ceramic packaging boxes of different specifications and perform auxiliary box annotation on the positions where ceramics are to be loaded. The establishment of the database aims to provide reliable training data and reference standards for subsequent detection models to achieve accurate identification of missing ceramics. First, different types of packaging box images are collected by an industrial camera in a standardized factory environment. To ensure data consistency, the shooting conditions are strictly controlled during the image acquisition process to minimize the image quality differences caused by light changes, shooting angle deviations, and reflection interferences. In addition, to ensure that subsequent algorithms can effectively extract the detailed information in the packaging box, the collected images will undergo preprocessing steps, including unifying the resolution, denoising, and color correction, to improve the clarity and stability of the images. Some pictures of empty packaging boxes are as shown in Figure 3 in (a).

[0072] Auxiliary box annotation. Use an artificial annotation tool to perform auxiliary box annotation on the area to be loaded inside the collected original packaging box photos. According to the internal design of each type of packaging box, precise auxiliary boxes are drawn for each reserved position for placing ceramics, as shown in Figure 3As shown in (b) and (c) thereof. These auxiliary box information are not only used to record the layout of different packaging boxes, but also provide key target position data during the training process of the model, so as to ensure that the detection system can efficiently identify the area to be loaded and reduce misjudgment caused by misalignment or occlusion. Finally, the labeled images and their auxiliary box information are classified and stored in the database according to the type and size of the packaging box, providing structured data support for the system. In practical applications, when a new packaging box image is collected, the system will retrieve similar packaging boxes from the database through feature matching and use the corresponding auxiliary box information to detect the loading status of the ceramics, so as to achieve efficient and accurate missing part identification.

[0073] Position matching. During the process of matching the packaging box samples with the database, ensuring the consistency of the captured images is the main challenge. The lighting conditions in the factory environment are unstable, coupled with equipment jitter and differences in shooting angles, resulting in significant changes in the images of the same packaging box taken at different times. To solve this problem, an autoencoder algorithm is introduced. The input image is compressed into a feature vector by the encoder and a high-quality image is reconstructed. The feature vector generated by the encoder is compared with the image feature vectors in the database through cosine similarity to identify the most matching image. After successful matching, the system applies the auxiliary box position information in the database image to the sample image to ensure the accuracy of ceramic missing detection. The decoder is responsible for reconstructing an image of the same size as the input image to ensure the output quality. The network structure of the autoencoder model in the embodiment of the present invention is as Figure 4 shown, and consists of two main parts: an encoder and a decoder. The encoder extracts features from the input packaging box image sample. Through a multi-layer convolutional neural network, key features are extracted layer by layer, and the high-dimensional image data is compressed into a low-dimensional feature vector, retaining important information in the image, such as shape, color, and texture, while removing redundant information. In this way, the packaging box image is represented as a 128-dimensional feature vector after being processed by the encoder. The decoder is responsible for reconstructing this feature vector into a reconstructed image close to the original image, converting the low-dimensional features into high-dimensional data through reverse operations, enabling the model to learn how to restore image details from the feature vector. In the position matching stage, after the autoencoder encodes a new sample into a feature vector, it is compared with the feature vectors in the database, and the closest database image is found by calculating the similarity. The associated auxiliary box position information is extracted to confirm the loading situation of the packaging box. This process ensures an accurate judgment of whether there is a missing part in the packaging box, thereby improving the overall efficiency and reliability of the production line.

[0074] The core task of the encoder is to convert the input image into a low-dimensional feature vector. In this stage, a 2D convolutional layer (Conv2D) is used to extract features from the image, and appropriate convolutional kernel parameters are set to generate feature maps. At the same time, LeakyReLU is used as the activation function to enhance the model's ability to capture important features and reduce the risk of neuron inactivation. In addition, the SENet (Squeeze-and-Excitation Network) channel attention mechanism is introduced. SENet compresses the feature map output by each convolutional layer into a channel descriptor through global average pooling to capture the global information of the feature map. Then, two fully connected layers (a compression layer and an expansion layer) are used to generate channel weights, and the calculation formula for the weights is:

[0075] ;

[0076] where z is the input feature map, and are the learned parameters, represents the LeakyReLU activation function, represents the sigmoid activation function. Finally, these weights are multiplied by the original feature map to complete feature recalibration, enhance the response of important features, and suppress less important features.

[0077] The main task of the decoder is to reconstruct the feature vector output by the encoder into the original image. Its structure is similar to that of the encoder. First, the feature vector is expanded into a larger feature map through a fully connected layer, and then deconvolution operations are used to convert the feature map into the same size as the input image. Through this process, the decoder ensures the quality of the reconstructed image, thereby guaranteeing the effectiveness of the subsequent matching process.

[0078] In the image retrieval stage, the cosine similarity is used to measure the similarity between the feature vectors of the sample image and the database image. Cosine similarity is an effective distance metric method that can evaluate the angle between two feature vectors, thereby reflecting their similarity degree. The calculation formula is:

[0079] ;

[0080] where A and B are two feature vectors. By calculating the cosine distance between the feature vectors, we can accurately find the most matching database image and extract the corresponding ceramic loading position information.

[0081] After completing the encoding and decoding operations, the Mean Square Error (MSE) is used to calculate the loss, and the weights of the autoencoder model are continuously adjusted by the optimizer to minimize the loss. MSE is used in the packaging box matching task to evaluate the pixel-level difference between the reconstructed image and the original image, ensuring high precision and consistency in the reconstruction process. By minimizing MSE, the model learns to capture the detailed features of the packaging box more accurately and optimizes the image feature extraction ability. The calculation formula is as follows:

[0082] ;

[0083] where n represents the number of pixel points, represents the actual value of the i-th, represents the predicted value of the i-th. During the optimization process, the Adam optimizer is used to adjust the model parameters to accelerate convergence and improve stability. Finally, by reducing the MSE loss, the model can accurately extract the features of the packaging box and improve the accuracy of matching.

[0084] The autoencoder model effectively extracts the key features of the packaging box image and combines cosine similarity for image matching, ensuring the accurate detection of the ceramic loading state. This method not only improves the efficiency and accuracy of detection but also provides reliable support for subsequent production processes, with high practical application value.

[0085] Specifically, stage two is the factory production inspection stage. In this stage, a product packaging box status detection model based on improved ResNet18 and CBAM is constructed to address multiple challenges in complex industrial environments; the overall framework of the product packaging box status detection model is as Figure 5 shown.

[0086] First, the input image is subjected to feature extraction through the CBR module composed of convolution, batch normalization (BN), and the ReLU activation function. Starting from low-level features such as edges and textures, feature mapping is performed through convolution operations, and BN is used to improve the training stability, while the ReLU activation function endows the model with non-linear capabilities. Next, downsampling is carried out through the MaxPool (max pooling layer) to reduce the data volume and retain the main features, significantly reducing the computational burden. Then, complex high-level features are further extracted through multiple improved residual blocks. The "skip connection" is introduced in the structure of the improved residual block, and this design can alleviate the vanishing gradient problem in deep networks, enabling information to flow smoothly between network layers and ensuring the stable transmission of deep features. The CBAM module is also added to the model to improve the accuracy of feature extraction. The CBAM module includes channel attention and spatial attention mechanisms, which can dynamically adjust the weights of important channels and regions in the feature map, making the model pay more attention to key features and ignore irrelevant information. This is particularly important for distinguishing objects with similar colors (such as ceramics and packaging boxes), enabling the model to accurately identify the target object even under lighting changes or occlusion conditions, significantly reducing the misjudgment rate.

[0087] In the final feature aggregation stage, the model introduces the Conv AvgPool module that combines convolution and average pooling. This module further extracts high-level features through convolution operations, while average pooling downsamples the feature map, retaining the overall information and reducing the feature dimension. This combined design ensures that key features are retained during the aggregation process, effectively reducing the computational complexity, enhancing the generalization ability and computational efficiency of the model. Finally, it is mapped to the output space through the fully connected layer, and the feature vector is converted into a probability distribution through the Softmax activation function, thereby outputting the classification results of "no ceramics" or "with ceramics", ensuring the accuracy and efficiency of packaging detection. When a missing loading situation is detected, the system will automatically issue an alarm, thus realizing the intelligence and stability of the production line. The design of the overall framework ensures the adaptability and robustness of the model, provides technical support for realizing intelligent packaging detection, and significantly improves the production efficiency.

[0088] The loss function of the product packaging box status detection model adopts the cross-entropy loss. The cross-entropy loss function is suitable for multi-classification tasks and can effectively measure the difference between the predicted probability distribution of the model and the true distribution. Specifically, the formula for cross-entropy loss can be expressed as:

[0089] ;

[0090] where the total number of samples is N, and the total number of categories is C. The true labels are represented using one-hot encoding, where indicates whether sample i belongs to category c. When sample i belongs to category c, ; otherwise . The probability that the model predicts that sample i belongs to class c is .

[0091] During the training process, the model gradually adjusts the parameters by minimizing the cross - entropy loss to improve the recognition accuracy for the two classes of "no ceramic" and "with ceramic". In addition, a regularization term is added to prevent overfitting to ensure that the model has good generalization ability in a complex industrial environment. The L2 regularization term can be expressed as:

[0092] ;

[0093] where is the regularization term, is the regularization coefficient, used to control the strength of regularization, are the model parameters.

[0094] Therefore, the final loss function of the product packaging box status detection model can be expressed as:

[0095] ;

[0096] Such a design aims to improve the classification performance and generalization ability of the model, making it perform excellently in a complex industrial environment.

[0097] Specifically, the product packaging box status detection model of the embodiment of the present invention adopts an improved ResNet18 network combined with a CBAM module. Among them, the improved ResNet18 network addresses the deficiency that the original residual block fails to fully utilize the role of the batch normalization layer, and introduces Adaptive Batch Normalization and the LeakyReLU activation function before the convolutional layer to form an improved residual block. The detailed description is as follows.

[0098] (1) The improved residual block includes a first improved residual layer and a second improved residual layer connected in sequence; the embodiments of the present invention include a first improved residual block (Res_a1 and Res_b1), a second improved residual block (Res_a2 and Res_b2), a third improved residual block (Res_a3 and Res_b3), and a fourth improved residual block (Res_a4 and Res_b4); among them, Res_a1, Res_a2, Res_a3, and Res_a4 are the first improved residual layers, adopting a dual-path interaction mechanism. The upper branch quickly extracts shallow texture features (such as the edge contour of the packaging box) through a single ALC, and the lower branch deeply excavates semantic information (such as the local gray change in the missing loading area) through two cascaded ALCs. Then, the dual-path features complete weighted integration, and finally, the ReLU activation module is used to constrain the non-linear range; Res_b1, Res_b2, Res_b3, and Res_b are the second improved residual layers, adopting a single-path structure. The input features are directly processed by two cascaded ALCs and then added to the skip connection. Its advantage lies in reducing the computational complexity.

[0099] Specifically, for the specific structure of the ALC, refer to Figure 6 As shown, ALC (Adaptive BN->LeakyReLU->Conv) includes Adaptive BN (variable batch normalization), LeakyReLU activation function, convolutional layer, Adaptive BN (variable batch normalization), LeakyReLU activation function, and convolutional layer connected in sequence, and skip connections are introduced at the input and output. The design of pre-positioning the normalization adjusts the dynamic distribution of the input features before convolution. Combining with the weak activation characteristic of the LeakyReLU activation function in the negative interval, it not only avoids the neuron death problem of the traditional ReLU but also enhances the adaptability to the sudden change of the industrial scenario data distribution through the adaptive parameter mechanism.

[0100] (2) The variable batch normalization dynamically adjusts the normalization parameters according to the features of different batches, improving the robustness of the model in scenarios of small batch training or data imbalance. This optimization further improves the scale distribution of the features, increases the gradient, accelerates the convergence of the learning process, effectively alleviates the gradient disappearance problem, and at the same time retains the identity mapping characteristic in the residual block. This improvement enhances the stability and adaptability of the model in the missing loading detection task. The core improvement of the variable batch normalization lies in dynamically adjusting the normalization parameters to adapt to the distribution characteristics of different batches of data. Traditional normalization normalizes each feature channel using fixed mean and variance, while the improved method adaptively adjusts the normalization parameters according to the feature statistics of the current batch by introducing a dynamic parameter adjustment mechanism. The specific implementation steps are as follows:

[0101] Dynamic mean and variance calculation: For each mini-batch of input feature maps, calculate their mean and variance, which is consistent with traditional BN.

[0102] Adaptive scaling and translation: After normalization, introduce learnable scaling factors ( ), and translation factors ( ), but their parameter updates no longer solely rely on backpropagation, but are dynamically adjusted in combination with the statistics of the current batch. For example, by introducing additional fully connected layers or attention mechanisms, dynamically generate and according to the batch features.

[0103] Always, assume the input feature is , the variable batch normalization is expressed as:

[0104] ; where and are the mean and variance of the current batch, and and are dynamically generated parameters.

[0105] The improved batch normalization enhances the robustness of the model in scenarios of mini-batch training or unbalanced data distribution, while alleviating the vanishing gradient problem and improving the stability of the feature distribution.

[0106] (3) The improved ResNet18 network uses the LeakyReLU activation function as the activation function in the improved residual block. In the standard ResNet-18 network, the ReLU (Rectified Linear Unit) activation function is used. However, when the input value is negative, ReLU will set its output to zero, which may cause some neurons to be in an inactive state for a long time, resulting in the phenomenon of neuron death. This phenomenon will hinder the weight update of the network and reduce the learning ability and performance of the model. To solve this problem, LeakyReLU is used as the activation function. LeakyReLU is an improved version of ReLU, inheriting the advantages of its simple and efficient calculation, and at the same time introducing a non-zero slope in the negative input region to avoid neuron death. Specifically, LeakyReLU outputs a certain proportion (usually 0.01 times) of its input value when the input value is negative, ensuring that there is still a non-zero gradient during backpropagation, so that the weights of the model can be continuously updated and the learning efficiency can be improved. Its mathematical definition is as follows.

[0107] The ReLU activation function is expressed as:

[0108] ;

[0109] The LeakyReLU activation function is expressed as:

[0110] ;

[0111] Among them, is a small constant (usually taken as 0.01).

[0112] The improvement of using the LeakyReLU activation function helps to alleviate the vanishing gradient problem, making the ResNet18 network more adaptable to deep learning tasks, thereby enhancing the ability to model brain tumor image features. In the missing component detection task, disturbances such as illumination changes and occlusions may cause some features to produce negative values. LeakyReLU allows the network to effectively capture edge features and fine-grained information in complex scenarios. At the same time, the introduced non-zero gradient improves the convergence speed and classification accuracy, enhances the robustness of the model, and provides a stable and reliable solution for intelligent detection.

[0113] (4) Introduce the CBAM module. The CBAM module aims to enhance the feature extraction ability of convolutional neural networks through the channel attention mechanism (Channel Attention, CAM) and the spatial attention mechanism (Spatial Attention, SAM). The channel attention mechanism filters important features in the channel dimension of the feature map, while the spatial attention mechanism assigns weights to spatial positions, thereby achieving effective fusion of global and local features. By enhancing the feature extraction ability of the model, the CBAM module improves the generalization ability of the detection model, enabling it to adapt to different packaging box layouts and diverse ceramic types. Combined with ResNet18, the introduction of CBAM effectively improves the robustness and detection accuracy of the model, ensuring that even in a complex industrial environment, the system can quickly identify missing components and issue alarms in a timely manner, guaranteeing the smooth progress of the production process. The attention mechanism of CBAM is as Figure 7 shown.

[0114] The channel attention mechanism assigns weights to different channels by extracting the global information of the feature map. Given the input feature map , first perform global average pooling and max pooling along the spatial dimension to obtain two feature vectors of size 1×1×C. Subsequently, these two feature vectors are processed by a shared two-layer multi-layer perceptron (MLP), and element-wise addition is performed. Finally, the channel attention weights are generated through the Sigmoid activation function. The calculation process is as follows:

[0115] ;

[0116] Among them represents the channel attention weight, is the Sigmoid function.

[0117] The spatial attention mechanism aims to highlight the important spatial regions in the feature map. For the input feature map , first, average pooling and max pooling are performed along the channel dimension to obtain two feature maps of size H×W×1. These two feature maps are concatenated along the channel dimension, and a convolution operation with a 7×7 convolutional kernel is used to generate the spatial attention weights. The specific calculation is as follows:

[0118] ;

[0119] where represents the spatial attention weight, represents the concatenation operation along the channel dimension.

[0120] Specifically, the packaging detection process of the S104 is as Figure 8 shown. The sample image of the sample to be detected is collected and matched with the database image to obtain the position information; according to the position information, the sample image is cut in a fixed area to extract the sub-image to be detected, and a series of image enhancement processes such as brightness, contrast, and noise suppression are performed to improve the image quality and recognizability. The enhanced sub-image will be input into the improved ResNet18 model for feature extraction. The improved ResNet18 effectively captures the multi-level features of the image through multiple layers of convolution and residual connections. The CBAM module focuses on the key features through the channel and spatial attention mechanisms, further improving the recognition ability of the model. After the feature extraction is completed, the model generates the classification probability distribution of each sub-image through the fully connected layer, and uses the Softmax activation function to convert the feature vector into a probability value, and finally determines whether the position is filled with ceramics. The detection result comprehensively judges the status of each sub-image as "no missing filling" or "missing filling" for the packaging box. If the system detects a missing filling situation, an alarm will be triggered and the missing filling position will be replenished, and then a sample detection process will be performed again to ensure the efficiency and stability of the production line.

[0121] To verify the effectiveness of the present invention, an experimental platform was built to implement the method of the present invention. The required equipment includes a hardware and software system: for the hardware part, an MV-CS200-10GC V5.0 camera (2000-megapixel IMX183 sensor) is used in conjunction with an MVL-KF1228M-12MPE lens (12-mm focal length), which is connected to an industrial computer through a gigabit Ethernet cable to ensure real-time acquisition of high-resolution images; the power supply system is provided with stable power support by a KPL-060M-VI adapter (24V / 2.5A) and a dedicated I / O cable; at the software level, the computing platform is equipped with an Intel Core i7-8750H CPU and an NVIDIA GeForce GTX1050Ti GPU (4GB video memory), which supports accelerated inference of deep learning models, and an improved ResNet18+CBAM model is built based on Python 3.8.5 and PyTorch 1.8.0. The test results are as Figure 9 shown. When the package is an empty box, the yellow light is on, indicating that no ceramics are loaded; when some ceramics are loaded in the packaging box, the red light is on, indicating a missing load; when all ceramics are loaded in the packaging box, the green light is on, indicating that the packaging is completed.

[0122] The present invention introduces a model that combines an improved ResNet18 and CBAM (Convolutional Block Attention Module) for ceramic packaging detection. In the ceramic packaging detection process, the system is divided into two main stages: the database establishment, auxiliary box annotation and matching stage, and the factory production sample detection stage. In the first stage, the main challenge faced by the system is to ensure the matching accuracy between the sample image and the database image. Due to the unstable lighting conditions, equipment jitter and shooting angle deviation in the factory environment, the images of the same packaging box taken at different times may have significant differences, which complicates the matching process. To solve this problem, this study uses an autoencoder algorithm, which consists of an encoder and a decoder. The encoder extracts key features through convolutional layers and uses the activation function LeakyReLU and the SENet channel attention mechanism to enhance the attention to important features. After the feature extraction is completed, the feature vector output by the encoder will be matched with the feature vector of the image in the database through cosine similarity to find the most similar database image. Once the matching is successful, the position information of the auxiliary box in the database image will be passed to the sample image to ensure the accurate detection of missing ceramics. This process not only improves the matching accuracy but also provides reliable position information for subsequent detection. In addition, the decoder is responsible for reconstructing the feature vector into a high-quality image of the same size as the input image, thus ensuring the effectiveness of the output. Through this efficient matching mechanism, the system can effectively overcome the influence caused by changes in shooting conditions and improve the overall accuracy and reliability of ceramic packaging detection.

[0123] During the factory production inspection and matching stage, the system faces multiple challenges such as the similar colors between the ceramics and the bottom of the packaging box, light changes, and object occlusion. These factors make it significantly difficult for the model to distinguish the ceramics from the background. First, according to the position information provided by the database, the sample images are cut in fixed areas to extract the sub-images of each position to be detected. Subsequently, these sub-images undergo a series of image enhancement processes to improve their quality and recognizability. The enhanced sub-images will be input into the improved ResNet18 and CBAM models for feature extraction. In feature extraction, the improved ResNet18 architecture effectively captures the multi-level features of the image through multiple layers of convolution and residual connections, generates high-dimensional feature maps, and enhances the ability to extract details. At the same time, the CBAM module further strengthens the model's focus on key features through channel and spatial attention mechanisms. The channel attention mechanism dynamically adjusts the weights of each feature channel to emphasize important features, and the spatial attention mechanism guides the model to focus on important regions in the image to ensure a high level of recognition accuracy under light changes and occlusion. After feature extraction, the model maps the high-dimensional features to the output space through a fully connected layer, generates the classification probability distribution of each sub-image, and applies the Softmax activation function to convert the feature vector into a probability value to finally determine whether ceramics are loaded at this position. To improve the accuracy and robustness of classification, the present invention uses the Cross-Entropy Loss function to optimize the classification effect of the model, ensuring more stable recognition of missed loading situations during the training process. The detection results are finally sorted into "no missed loading" or "there is missed loading" to ensure the efficiency and stability of the production line. If the system determines a missed loading, it will automatically trigger an alarm and refill the missed loading position, and then re-execute the sample detection process. Through this systematic feature extraction and classification process, the model of this study demonstrates excellent adaptability and robustness in complex industrial environments, laying a solid foundation for the accuracy and reliability of intelligent packaging detection.

[0124] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A ceramic packaging detection method based on an autoencoder and an improved ResNet18 model, characterized in that, It includes the following steps: Database construction step, constructing a ceramic packaging box image database; Model construction step, constructing a product packaging box status detection model by combining the improved ResNet18 and CBAM; Image acquisition step, acquiring an image of the ceramic packaging box to be detected as a detection image, and using an autoencoder model to match the detection image with the ceramic packaging box image database to obtain position information; Image segmentation step, using the position information to perform regional segmentation on the detection image and extract the sub-images to be detected; Packaging detection step, inputting the sub-images into the product packaging box status detection model to detect whether there is a shortage of ceramics in the packaging; if there is no shortage, the detection ends, and if there is a shortage, the ceramics are replenished and the image acquisition step is re-entered; The product packaging box status detection model includes a CBR module, a max pooling layer, a first improved residual unit, a first CBAM module, a second improved residual unit, a second CBAM module, a Conv AvgPool module, and a fully connected layer connected in sequence; the CBR module consists of convolution, batch normalization, and a ReLU activation function, and is used to extract features from the input image; the max pooling layer is used to downsample the features extracted by the CBR module; The first improved residual unit includes an improved residual block, which is used to further extract features from the features output by the max pooling layer; The first CBAM module includes a channel attention and a spatial attention mechanism, which are used to dynamically adjust the weights of important channels and regions in the feature map output by the first improved residual unit; The second improved residual unit includes several improved residual blocks, which are used to perform multiple further feature extractions on the feature map output by the first CBAM module; The second CBAM module includes a channel attention and a spatial attention mechanism, which are used to dynamically adjust the weights of important channels and regions in the feature map output by the second improved residual unit; the Conv AvgPool module further extracts features from the feature map output by the second CBAM module through convolution operations and downsamples the extracted features through average pooling; the fully connected layer converts the features output by the Conv AvgPool module into a probability distribution, thereby outputting a classification result of whether there is a shortage of ceramics in the packaging; The improved residual block includes a first residual layer and a second residual layer connected in sequence. The first residual layer adopts a dual-path interaction mechanism. One path quickly extracts shallow texture features through a single ALC, and the other path deeply mines semantic information through two cascaded ALCs. Finally, the extraction results of the two paths are added and then passed through the ReLU activation function; the second residual layer adopts a single-path structure. The input features are directly processed by two cascaded ALCs and then added to the input features through a skip connection, and then passed through the ReLU activation function; The ALC includes a variable batch normalization, a LeakyReLU activation function, a first convolutional layer, a variable batch normalization, a LeakyReLU activation function, and a second convolutional layer connected in sequence. The input of the ALC is added to the output of the second convolutional layer through a skip connection to obtain the final output of the ALC; The variable batch normalization is expressed as: where x represents the original input, represents the result of variable batch normalization; μ B and σ B are the mean and variance of the current batch, γ adaptive and β adaptive are dynamically generated parameters; ∈ is a positive number used to prevent division by zero.

2. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 1, wherein The construction of the ceramic packaging box image database includes the following steps: Collect images of ceramic packaging boxes without ceramics installed; Manually annotate the collected images, mark the empty spaces in the packaging boxes where ceramics need to be installed, and perform auxiliary box annotation.

3. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 1, wherein The CBAM module performs the following steps on the input feature map: For the input feature map F, calculate the channel attention weight M c (F) and the spatial attention weight M s (F), respectively, which are expressed as: M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); M s (F) = σ(Conv 7×7 [AvgPool(F), MaxPool(F)]); Among them, σ is the Sigmoid function, [] represents the concatenation operation in the channel dimension; MLP represents the multi-layer perceptron operation; AvgPool represents the average pooling operation; MaxPool represents the maximum pooling operation; Conv 7×7 represents the convolution operation using a 7×7 convolution kernel; Multiply the channel attention weight by the original feature map to obtain the channel-weighted feature map F * , which is expressed as: Among them, represents element-wise multiplication; Multiply the spatial attention weight with the feature map F after channel weighting * to obtain the final output feature map F ** , which is expressed as:

4. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 1, characterized in that The loss function of the product packaging box status detection model is expressed as: Among them, L total represents the loss function of the product packaging box status detection model; represents the cross-entropy loss. The total number of samples is N, the total number of categories is C, and y i,c indicates whether sample i belongs to category c. If so, y i,c = 1, otherwise y i,c = 0; is the predicted probability that sample i belongs to category c; R(θ) is the regularization term, λ is the regularization coefficient used to control the strength of regularization, M represents the total number of all trainable parameters in the model, j represents the parameter index, and θ j are the model parameters.

5. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 1, wherein The use of the autoencoder to match the detection image with the ceramic packaging box image database includes the following steps: Use the encoder to compress the input detection image into a feature vector and reconstruct a high-quality image; Compare the feature vector generated by the encoder with the images in the ceramic packaging box image database, identify the most matching image, and extract the position information therein; The decoder uses the feature vector output by the encoder to reconstruct an output image of the same size as the input detection image.

6. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 5, wherein, The use of the encoder to compress the input detection image into a feature vector specifically uses a 2D convolutional layer to extract features from the detection image, and introduces the SENet channel attention mechanism to calculate the channel weights during feature extraction, which is expressed as: s = σ(W2δ(W1z)); Where z represents the original feature map, W1 and W2 are learned parameters, δ represents the LeakyReLU activation function, and σ represents the sigmoid activation function; Multiply the weight s by the original feature map to complete feature recalibration.

7. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 5, characterized in that, The decoder uses the feature vector output by the encoder to reconstruct an output image of the same size as the input detection image. Specifically: expand the feature vector into a larger feature map through a fully connected layer, and then use a transposed convolution operation to convert the feature map into the same size as the input image.

8. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 6, characterized in that The comparison of the feature vector generated by the encoder with the images in the ceramic packaging box image database is specifically done by calculating the cosine similarity between the feature vector generated by the encoder and the image feature vectors in the ceramic packaging box image database. The calculation formula of the cosine similarity is: Where A and B are two feature vectors.

9. The ceramic packaging detection method based on the autoencoder and the improved ResNet18 model according to claim 1, characterized in that The loss function of the autoencoder is: Among them, n represents the number of pixel points, and Y i represents the actual value of the i-th pixel point, and y i represents the predicted value of the i-th pixel point output by the decoder; MSE represents the mean square error.

Citation Information

Patent Citations

  • Cigarette package anomaly detection and positioning method based on deep learning

    CN111951264A

  • Image Authentication Method and Real-Time Product Authentication System

    US20210397897A1