Pet food packaging bag image registration detection method
Through the multimodal image registration detection method, combined with RGB and thermal imaging images, deep learning models are used to perform image registration, segmentation and defect detection, which solves the problem of low accuracy and automation level of printing quality detection in the prior art, and achieves high-precision and high automation detection effects.
Patent Information
- Application Number
- CN202510122121.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art has limitations in single-modal image processing, insufficient application of multimodal fusion technology, adaptability of deep learning technology, and low automation level in the printing quality detection of pet food packaging bags, resulting in low detection accuracy and robustness.
Using the multimodal image registration detection method, images are collected through RGB cameras and thermal imaging cameras, standardized preprocessing and interpolated to the same resolution, input the trained image registration model, and output the printing quality detection results. The image registration model includes an image registration module, an image segmentation module and a defect detection module. It uses deep learning models such as UNet and YOLOv5 for feature extraction, segmentation and defect detection.
It realizes high accuracy and high automation of printing quality inspection of pet food packaging bags, improves the robustness and accuracy of inspection, especially in complex backgrounds and diverse defect detection.
Smart Images

Figure CN120031841A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image analysis based on computer vision, and in particular relates to a pet food packaging bag image registration detection method. Background Art
[0002] As an important form of food packaging, pet food packaging bags not only play a key role in protecting food safety and extending shelf life, but also attract consumers to buy through exquisite printing designs. However, with the advancement of packaging industry technology and the increase in market demand, the printing quality of packaging bags directly affects the brand image and product competitiveness. High-quality packaging bags usually have complex patterns, multiple colors and precise text layout, so the inspection of printing quality has become a crucial part of the production process.
[0003] Traditional printing quality inspection methods mostly rely on manual visual inspection or simple machine vision systems. Although manual inspection is flexible, it is inefficient and easily affected by subjective factors; while traditional machine vision inspection methods are mostly based on single-mode image processing technology, such as RGB image analysis. These methods show obvious limitations when dealing with complex printed patterns, especially when facing printed designs with similar colors, complex textures or fine details, the detection accuracy and robustness are low.
[0004] In recent years, the rise of multimodal image processing technology has provided new solutions for printing quality inspection. By combining RGB images and thermal imaging images, multidimensional information of the surface of packaging bags can be obtained, including color, texture and temperature distribution. This multimodal fusion technology can effectively make up for the shortcomings of a single mode, especially in the detection of material defects, printing offset and color unevenness. In addition, the rapid development of deep learning technology has provided powerful algorithmic support for image registration and defect detection. Based on convolutional neural networks (CNN) and target detection models (such as YOLO and Faster R-CNN), automated and high-precision detection of packaging bag printing quality can be achieved.
[0005] Although multimodal fusion technology and deep learning algorithms have been widely used in the field of industrial inspection, the inspection of the printing quality of pet food packaging bags still faces many challenges. This is mainly because pet food packaging bags are usually made of flexible materials, which are prone to wrinkles, reflections or temperature changes on the surface, which further increases the complexity of inspection. Existing technical attempts have the following defects:
[0006] 1. Limitations of single-modal image processing: Traditional printing quality inspection systems are mostly based on RGB image analysis, which has insufficient resolution for areas with similar colors. RGB images are difficult to distinguish for patterns or texts with similar colors, which can easily lead to missed detection or false detection. At the same time, they are highly dependent on lighting conditions, and changes in ambient lighting can significantly affect the detection effect, especially in highly reflective or low-contrast scenes. They lack the ability to detect material defects, and RGB images cannot provide the physical properties of materials (such as temperature distribution), so it is difficult to detect hidden defects on the surface of materials.
[0007] 2. Insufficient application of multimodal fusion technology: Although multimodal technology has made significant progress in the fields of medical imaging and remote sensing image processing, its application in packaging and printing quality inspection is still immature. The viewing angles and resolutions of multimodal images (such as RGB and thermal imaging) vary greatly, and traditional registration methods based on geometric features are difficult to meet high-precision requirements. Existing methods mostly use simple pixel-level superposition and lack deep fusion of multimodal features, resulting in low information utilization.
[0008] 3. Adaptability issues of deep learning technology: Deep learning performs well in the field of image detection, but its application in printing quality inspection still faces the following challenges. The lack of public, standardized multimodal datasets in the field of packaging bag printing quality inspection limits the training and optimization of deep learning models. Deep learning models are sensitive to wrinkles, reflections, and deformations on the surface of packaging bags, which can easily lead to unstable detection results. Industrial inspection requires high precision and real-time performance, but deep learning models are usually computationally complex and difficult to deploy directly to edge devices.
[0009] 4. Low automation level of existing systems: Traditional printing quality inspection systems often require manual assistance and have a low degree of automation. In the image annotation, feature extraction and defect classification process, a large amount of manual participation is required, which is inefficient. It is difficult to feed back the inspection results to the production line in real time, and it is impossible to automatically remove or repair defects.
[0010] Therefore, designing an efficient and accurate printing quality inspection method based on multimodal image registration has important practical significance. Summary of the invention
[0011] In view of the above problems, the present invention provides a pet food packaging bag image registration detection method, which is characterized by comprising the following steps:
[0012] S1, collects RGB images and thermal imaging images of pet food packaging bags through RGB camera and thermal imaging camera respectively;
[0013] S2, standardizes and preprocesses the RGB image and thermal image of the pet food packaging bag, and interpolates the thermal image to the same resolution as the RGB image;
[0014] S3, inputting the pre-processed pet food packaging bag RGB image and thermal imaging image into the trained image registration model, and outputting the pet food packaging bag printing quality inspection result;
[0015] The image registration model includes an image registration module, an image segmentation module and a defect detection module;
[0016] The image registration module performs feature extraction, feature point matching, matching point pair optimization and fusion on the collected RGB image and thermal imaging image to generate a pixel-level fused image;
[0017] The image segmentation module uses UNet as a segmentation model to segment the packaging bag area and the printed pattern area from the input pixel-level fusion image. The UNet consists of an encoder and a decoder. The encoder is responsible for extracting features and gradually reducing the resolution, while the decoder is responsible for restoring the spatial resolution of the image.
[0018] The defect detection module performs detection based on the defect area image obtained after segmentation, and outputs the position information and classification result of each printing quality defect.
[0019] Preferably, the process of constructing the data set for training the image registration model is:
[0020] Collect samples of pet food packaging bags from different batches, including normal samples and defective samples; the types of defects include missing printing, offset printing, and blurred printing;
[0021] The images collected by the RGB camera and the thermal imaging camera have the same viewing angle and uniform brightness. The acquisition frequency and the acquisition time of each sample are designed to ensure that the complete area of the packaging bag is covered;
[0022] The printing defects in the RGB image and the thermal imaging image are annotated, and the defects in the printed part of the pet food packaging in the image are annotated with a rectangular box, and the rectangular bounding box is B = {x min ,y min , x max ,y max},(x min ,y min ) and (x max ,y max ) represent the coordinates of the upper left corner and the lower right corner of the bounding box respectively; defects are divided into three types: printing blur, printing offset and printing missing, which are represented by codes 0, 1 and 2 respectively.
[0023] Preferably, the feature extraction in the image registration module is to perform multi-layer feature extraction, wherein each layer uses a pre-trained convolutional neural network to extract high-dimensional features of RGB and thermal imaging images:
[0024] F RGB =f CNN (I RGB )
[0025] F T =f CNN (I T1 )
[0026] Among them, f CNN represents the convolutional neural network, F RGB and F T are the feature maps of RGB images and thermal images respectively; the dimension of the output feature map is F∈R H′×W ' ×C , where H′ and W′ represent the height and width of the feature map, which is the downsampling of the input image resolution; C represents the number of feature channels, which represents the extracted feature dimension;
[0027] The pre-trained convolutional neural network includes a convolution layer, a pooling layer, and a normalization layer;
[0028] The convolutional layer is used to extract local features:
[0029]
[0030] Among them, F l (x, y, c) represents the value of channel c at position (x, y) in the output feature map of the lth layer, i.e., the output feature map of the lth layer; W l (i, j, c) represents the convolution kernel weight, and the kernel size is (2k+1)×(2k+1); b l is the bias term, k is the convolution kernel radius;
[0031] The pooling layer uses the maximum pooling operation to downsample and reduce the spatial resolution of the feature map:
[0032]
[0033] Where R represents the pooling window range, x′, y′ represent the coordinates after pooling;
[0034] Finally, batch normalization and nonlinear activation function ReLU are used to enhance the feature expression ability and obtain a multi-layer feature map F RGB and F T .
[0035] Preferably, the feature point matching in the image registration module is to achieve cross-modal key point alignment through feature point detection and matching, and establish a corresponding relationship between multi-modal images:
[0036] First, key point extraction is performed; key points P are detected in each feature map:
[0037] P = {p i |i=1,2,...,N}
[0038] Wherein, each key point p includes two-dimensional coordinates (x, y) and feature descriptor d, and N is the number of key points detected in each feature map;
[0039] Next, a feature descriptor d is generated for each key point, extracted from its local feature map:
[0040] d(p)=desc(F,p)
[0041] Where desc(F, p) is expressed as RGB , F T} extract the local feature descriptor with p as the center;
[0042] Secondly, generate matching point pairs; use the similarity measure between feature descriptors to match key points in different feature maps:
[0043]
[0044] Among them, sim(d 1 , d 2 ) represents the similarity measurement function. The present invention adopts Euclidean distance calculation, τ is the similarity threshold, is the i-th pixel in the RGB image, is the jth pixel in the thermal imaging image.
[0045] Preferably, the matching point pair optimization and fusion in the image registration module firstly uses the random sampling consensus algorithm RANSAC to remove the wrong matching point pairs and retain the internal point set. f The set of matched point pairs after optimization is:
[0046] M f =RANSAC(M)
[0047] Then, the homography transformation matrix is estimated based on the optimized matching point pair M f , estimate the homography transformation matrix H from thermal imaging to RGB image:
[0048]
[0049] Among them, h ** Respectively represent the effects of transformation and perspective transformation in the x direction and y direction; the homography transformation relationship is:
[0050]
[0051] Where (x, y) and (x′, y′) are the coordinates of the matching points in the thermal image and RGB image respectively;
[0052] Finally, pixel-level image fusion is performed. During the fusion process, the pixel values of multiple images are merged into a final image. Set the weight coefficient α, according to the weighted average method:
[0053] I R =αI RGB +(1-α)I T
[0054] Finally, I R is the registered image obtained after image registration processing.
[0055] Preferably, the image segmentation module adopts a UNet model, wherein the encoder consists of a series of convolutional layers and pooling layers;
[0056] At each layer, the input registered image is convolved, and nonlinearity is introduced through the activation function ReLU. R As the input of layer 0, it undergoes the convolution operation:
[0057]
[0058] Represents the output feature map after the convolution operation; is the parameter of the convolution kernel;
[0059] Next, the pooling operation is performed. The pooling operation is used for downsampling, dynamic context modeling is introduced, and the adaptive pooling operation generates multi-scale context features:
[0060]
[0061] The decoder aims to restore the low-resolution feature map to the size of the original image and combine the high-resolution features of the encoder part through skip connections;
[0062] The multi-scale context feature maps are mapped back to the original resolution, and the feature maps are upsampled by deconvolution operations through weighted fusion:
[0063]
[0064] is the upsampled feature map, the size is gradually restored, l represents the lth layer; in the decoder, the jump connection splices the high-resolution features of the encoder with the upsampled results of the decoder to retain the details:
[0065]
[0066] in, is the concatenated feature map, is a learnable weight parameter used to dynamically adjust the contribution of contexts of different scales;
[0067] The decoder convolves the concatenated feature map and performs nonlinear transformation through ReLU activation:
[0068]
[0069] Finally, the decoder output F fin The number of channels is reduced to 1 through a 1×1 convolution layer, and the category probability of each pixel is mapped to the interval [0, 1] through the Sigmoid activation function, so that each pixel belongs to the bag area R p and the printed pattern area R s Probability
[0070]
[0071] in, is the output probability map, with a value of [0, 1], indicating the probability that each pixel belongs to its respective category, σ is the activation function, and k represents the kth image;
[0072] Next, by transforming the probability map Binarization to obtain the packaging bag area and printed pattern area Using threshold when Then the image (x, y) belongs to the packaging bag area, otherwise, it belongs to the printing area, and finally the segmented image I is obtained. seg .
[0073] Preferably, the UNet model uses a cross entropy loss function to train the model, the goal being to minimize the difference between the predicted result and the true label. The formula of the cross entropy loss function is as follows:
[0074]
[0075] Among them, Y k (x, y) represents the true label, indicating whether the pixel (x, y) belongs to the target area. H and W are the height and width of the segmentation map, respectively. Through back propagation, the parameters of the model are updated to minimize the loss function.
[0076] Preferably, the defect detection module includes two parts: a dynamic feature enhancement unit that introduces modality guidance and a YOLOv5 target detection unit;
[0077] The dynamic feature enhancement unit performs channel compression and global modeling on the thermal imaging features to generate guidance weights:
[0078] w T =σ(FC(Gap(F T )))
[0079] Among them, Gap(·) represents global average pooling, which is used to extract global features, FC(·) represents the fully connected layer, which maps the global features to the weight space, and σ represents the Sigmoid activation function;
[0080] Use thermal imaging weights to adjust the channel importance of RGB features for dynamic feature enhancement:
[0081]
[0082] The enhanced RGB features are fused with the original thermal imaging features, and the convolution operation is used to further extract the local information in the fused features:
[0083]
[0084] Using the YOLOv5 target detection unit and the enhanced feature F en Enables defect detection in pet food packaging bags.
[0085] Preferably, the YOLOv5 target detection unit and the enhanced feature F are used. en The specific process of realizing defect detection of pet food packaging bags is as follows:
[0086] The segmented image I seg Input into the YOLOV5 target detection unit for grid division and bounding box prediction; for the kth input image First, it is divided into s×s grids; then the bounding box prediction is performed for each grid unit. For the i-th network unit gi, the YOLOV5 target detection unit outputs the predicted bounding box B of the network unit i =(x i ,y i , w i ,h i ), (x i ,y i , w i ,h i ) is the position coordinate of the predicted bounding box; for the i-th bounding box, YOLOv5 will output a category probability C i , represents the probability that the bounding box belongs to a certain type of defect;
[0087] After the target detection is completed, a set of bounding boxes and category probabilities are generated. According to the detection results, non-maximum suppression NMS is set. For multiple overlapping bounding boxes, NMS will be based on the confidence threshold δ NMS Filter out the best bounding boxes and remove boxes with high overlap and low confidence;
[0088] For the bounding box B i =(x i ,y i , w i ,h i ) and B j , calculate the intersection of two bounding boxes:
[0089]
[0090] If the IoU of two bounding boxes is greater than the threshold δ NMS , then the bounding box with higher confidence is retained and the bounding box with lower confidence is removed;
[0091] Adopt F en Calculate the confidence of the printing defects in the grid unit to correct the detection result of YOLOV5; for the i-th network unit g i Defective confidence Use the feature pyramid network FPN to fuse features of different resolutions. en Input into the multi-scale feature layer of the pyramid network:
[0092]
[0093] in, is the original confidence value predicted by the model, is the normalized confidence in the interval [0, 1] obtained by the Sigmoid activation function;
[0094] The defect category probability output by YOLOv5 and the confidence output by the multi-scale feature layer of the pyramid network are combined and input into the softmax activation function to obtain the printing defect detection category class; for the i-th network unit g i , the specific calculation formula is as follows:
[0095]
[0096] Defects are divided into three categories: printing blur 0, printing offset 1, and printing missing 2. If class i = 0 means that the area selected by the predicted bounding box in the i-th grid unit has a printing blur problem. i= 1 means that the area selected by the predicted bounding box in the i-th grid unit has a printing offset problem. i =2 means that there is a printing missing problem in the area selected by the predicted bounding box in the i-th grid unit.
[0097] Compared with the prior art, the present invention has the following beneficial effects:
[0098] 1. Combination of multimodal image registration and deep learning: The present invention proposes a detection method combining multimodal image registration and deep learning algorithm for the detection of printing quality of pet food packaging bags. By jointly using RGB images and thermal imaging images, multidimensional information of the packaging bag is obtained, and high-precision alignment is achieved through deep learning-assisted registration. RGB images provide color and texture information of the packaging bag, and thermal imaging images capture the temperature distribution of the material for detecting hidden defects. Convolutional neural networks (CNNs) are used to extract high-dimensional features of multimodal images, and high-precision alignment of multimodal images is achieved by combining feature point matching and homography transformation matrix estimation. The data set is expanded by rotation, scaling, and noise addition, and random sampling consistency is used to optimize the feature point matching results to improve the robustness and accuracy of the registration.
[0099] 2. Dynamic context-aware module improves segmentation accuracy: This invention innovatively introduces a dynamic context-aware module to address the problem of insufficient segmentation accuracy of the traditional UNet model under complex backgrounds. Multi-scale context features are generated through adaptive pooling to capture the global information of the packaging bag background. The contribution of context features of different scales is dynamically adjusted using learnable weights to improve the adaptability to complex backgrounds. The high-resolution features of the encoder and the low-resolution features of the decoder are combined to retain edge details through jump connections to improve edge blur.
[0100] 3. Modality-guided dynamic feature enhancement module: The present invention designs a modality-guided dynamic feature enhancement module, which is combined with the YOLOv5 target detection framework to improve the detection capability of printing defects. The guided weights are generated by thermal imaging features, and the channel importance of RGB features is dynamically adjusted to enhance the model's perception of defective areas. The enhanced RGB features and the original thermal imaging features are integrated to generate a more robust multimodal feature map. The modality-guided dynamic feature module is seamlessly integrated into the YOLOv5 framework, which significantly improves the detection accuracy of complex defects (such as color unevenness, printing offset, and blur) while maintaining real-time detection capabilities.
[0101] 4. Evaluation of the integrity of text and patterns on packaging bags: The present invention further designs a text and pattern integrity evaluation module based on defect detection results, and innovatively proposes the following detection method: By comparing the color distribution of the pattern area of the packaging bag with the standard template, the color difference is calculated to detect color inconsistency defects. The defect area marked by the defect detection model is used to determine the fuzzy area and the degree of text offset in combination with the template information. By analyzing the area marked as missing, the missing text or pattern on the packaging bag can be accurately identified. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] Figure 1 It is an overall flow chart of the image registration detection method of the present invention.
[0103] Figure 2 Data collection and processing flow chart of the present invention.
[0104] Figure 3 Data processing flow chart of the image registration module of the present invention.
[0105] Figure 4 Data processing flow chart of the image segmentation module and defect detection module of the present invention.
[0106] Figure 5 This is a comparison chart of the image registration printing quality detection performance of the pet food packaging bag of the present invention. DETAILED DESCRIPTION
[0107] The overall logic of the solution of the present invention is as follows Figure 1 As shown, the invention is further described below in conjunction with specific embodiments.
[0108] 1. Data Collection
[0109] The purpose of this stage is to design and implement the system architecture, data acquisition and annotation, camera calibration and image preprocessing process based on multimodal imaging, and provide a hardware foundation for the printing quality inspection of pet food packaging bags, such as Figure 2 As shown in the figure. By using RGB cameras and thermal imaging cameras together, multi-dimensional information of packaging bags is obtained, which improves the comprehensiveness and accuracy of detection. During data collection, standardized imaging settings and annotation processes are used to ensure the diversity of samples and the accuracy of ground truth. The multimodal fusion images generated by image preprocessing provide high-quality data input for subsequent deep learning detection, ensuring the detection effect and robustness of the entire system.
[0110] 1. System architecture design
[0111] The present invention uses an RGB camera, a thermal imaging camera and a light source system. The RGB camera collects the color and texture information of the packaging bag and generates an RGB image. RGB; The thermal imaging camera is used to capture the temperature distribution on the packaging bag surface, detect hidden defects, and generate thermal imaging images. T ; The light source system uses a uniform and adjustable light source to ensure image quality and reduce reflection interference. Install them on the mounting platform, fix the camera and light source, ensure consistent imaging angle, and place the packaging bag in a fixed position. Secondly, it is equipped with a GPU computing unit to support the training and reasoning of deep learning models.
[0112] 2. Data collection and annotation
[0113] Collect samples of pet food packaging bags from different batches, including normal samples S n and defective sample S a The types of defects include missing prints, offset prints, and blurred prints. The RGB camera and thermal imaging camera are fixedly installed to ensure that the viewing angle of the collected images is consistent, and the imaging distance D and the light source intensity of the light source system are set to I light , to ensure uniform image brightness. RGB image I RGB and thermal imaging image I T Synchronous acquisition ensures that each pair of images corresponds to each other.
[0114] The acquisition frequency of the camera in the present invention is 5 frames per second, and the acquisition time of each sample is 2 seconds. It is ensured that images under sufficient angles and lighting conditions are collected to cover the complete area of the packaging bag. There are 210 normal samples and 130 defective samples, and a total of 3400 images are obtained.
[0115] Furthermore, the data annotation tool labellmg is used to annotate the printing defects in the RGB image and the thermal imaging image. Specifically, the defects in the printed part of the pet food packaging in the image are annotated with a rectangular box, and the rectangular bounding box is B = {x min ,y min , x max ,y max},(x min ,y min ) and (x max ,y max ) represent the coordinates of the upper left corner and lower right corner of the bounding box respectively. Defects are divided into three types: printing blur, printing offset and printing missing, which are represented by codes 0, 1 and 2 respectively.
[0116] Furthermore, the collected images are enhanced to increase data diversity. The input of data enhancement is the original multimodal image I RGB and I T , which together constitute the dataset I o .
[0117] First, the image is rotated. The image is rotated by an angle θ∈[-θmax ,θ max ],θ max is the maximum rotation angle:
[0118] I 1 (x, y) = I o (xcosθ-ysinθ, xsinθ+ycosθ)
[0119] Among them, (x, y) represents the coordinates of the pixel point.
[0120] Secondly, the image is scaled by the scaling factor s∈[s min ,s max ], scaled according to the following formula:
[0121] I 2 (x, y) = I o (s*x,s*y)
[0122] Again, noise is added. Add Gaussian noise N(0, σ 2 ), where σ is the standard deviation of the noise:
[0123] I 3 (x, y) = I o (x, y)+δ,δ~N(0,σ 2 )
[0124] The image I obtained after rotation, scaling and noise addition 1 , I 2 , I 3 Merge with the original image dataset to form the final image dataset I f In the process of image data enhancement, new images are generated, in which RGB images are merged into the RGB subset, and thermal images are merged into the thermal imaging subset.
[0125] 3. Data interpolation:
[0126] Dataset I f It is composed of the original image and the image after data enhancement:
[0127] I f = {I RGB , I T}
[0128] The thermal imaging image is interpolated to the same resolution H×W as the RGB image, where H is the image height and W is the image width.
[0129] I T1 = r(I T , H, W)
[0130] Where r is the interpolation function.
[0131] 2. Image Registration Module Design
[0132] The purpose of image registration is to convert the acquired RGB image I RGB and thermal imaging image I T Accurately align and generate multimodal fusion image T RGB-T The present invention uses deep learning-assisted registration technology to perform image registration on pet food packaging bags. f Perform feature extraction, feature point matching, transformation matrix estimation and optimization to achieve RGB image I RGB and thermal imaging image I T The overall process is as follows: Figure 3 shown.
[0133] 1. Feature extraction
[0134] Feature extraction is the core step of image registration, and its purpose is to extract f High-dimensional features are extracted from RGB images and thermal images in order to achieve key point matching and image alignment. In the application scenario of pet food packaging bags, feature extraction needs to consider the complex patterns, text layout, diverse colors and possible thermal imaging characteristics of the packaging bags.
[0135] First, from the dataset I f Select a pair of RGB and thermal imaging images (I RGB , I T1 ). Perform feature extraction on each pair of images separately.
[0136] Use pre-trained convolutional neural networks to extract high-dimensional features of RGB and thermal imaging images. Feature extraction process:
[0137] F RGB =f CNN (I RGB )
[0138] F T =f CNN (I T1 )
[0139] Among them, f CNN represents a convolutional neural network. RGB and F T They are the feature maps of RGB images and thermal images respectively. The dimension of the output feature map is F∈R H′×W′×C , where H′ and W′ represent the height and width of the feature map, which is the downsampling of the input image resolution. C represents the number of feature channels, which represents the extracted feature dimension, and is set to 256 in the present invention.
[0140] Secondly, the design of the pre-trained convolutional neural network includes the following: Designing the convolution layer. The convolution operation is used to extract local features:
[0141]
[0142] Among them, F l (x, y, c) represents the value of channel c at position (x, y) in the output feature map of the lth layer, that is, the output feature map of the lth layer. l (i, j, c) represents the convolution kernel weight, and the kernel size is (2k+1)×(2k+1). l is the bias term, and k is the convolution kernel radius.
[0143] Next, design the pooling layer. Use the maximum pooling operation to downsample and reduce the spatial resolution of the feature map:
[0144]
[0145] Among them, R represents the pooling window range, and x′, y′ represent the coordinates after pooling.
[0146] Finally, normalization is performed. Batch normalization and nonlinear activation function ReLU are used to enhance feature expression capabilities. The features of each channel are normalized:
[0147]
[0148] F ReLU =max(0, F BN )
[0149] Among them, μ and σ are the mean and standard deviation of the feature map respectively.
[0150] Furthermore, multi-layer feature extraction. The feature extraction network consists of multi-layer convolution conv, pooling layer Pool, normalization BN and activation function ReLU stacking. For the input F of the lth layer l-1 and output F l , the recursive relation is:
[0151] F l =ReLU(BN(Pool(conv(F l-1 ))))
[0152] 2. Feature point matching unit
[0153] In the feature extraction stage, multi-layer feature maps F are extracted from RGB images and thermal imaging images respectively after being processed by convolutional neural networks. RGB , F T , the total feature map F∈{F RGB , FT These feature maps contain high-dimensional feature information of multimodal images. Next, we need to achieve cross-modal key point alignment through feature point detection and matching. Establish the correspondence between multimodal images.
[0154] First, perform key point extraction. Detect key points P in each feature map:
[0155] P = {p i |i=1,2,...,N}
[0156] Each key point p includes two-dimensional coordinates (x, y) and a feature descriptor d, and N is the number of key points detected in each feature map.
[0157] Next, a feature descriptor d is generated for each key point, extracted from its local feature map:
[0158] d(p)=desc(F,p)
[0159] Where desc(F, p) is expressed as RGB , F T} extracts the local feature descriptor centered on p.
[0160] Secondly, generate matching point pairs. Use the similarity measure between feature descriptors to match key points in different feature maps:
[0161]
[0162] Among them, sim(d 1 , d 2 ) represents the similarity measurement function. The present invention adopts Euclidean distance calculation, τ is the similarity threshold, is the i-th pixel in the RGB image, is the jth pixel in the thermal imaging image.
[0163] 3. Matching point pair optimization
[0164] After the matching points are generated, the matching point pairs are optimized to remove the wrong matching points. The present invention uses the random sampling consensus algorithm (RANSAC) to remove the wrong matching point pairs. f The set of matched point pairs after optimization is:
[0165] M f =RANSAC(M)
[0166] Furthermore, the homography transformation matrix is estimated. Based on the optimized matching point pair M f , estimate the homography transformation matrix H from thermal imaging to RGB image:
[0167]
[0168] Among them, h ** They represent the effects of transformation and perspective transformation in the x and y directions respectively. The homography transformation relationship is:
[0169]
[0170] Among them, (x, y) and (x′, y′) are the coordinates of the matching points in the thermal imaging image and the RGB image, respectively.
[0171] Furthermore, pixel-level image fusion is performed. During the fusion process, the pixel values of multiple images are merged into a final image. Assume the weight coefficient α, according to the weighted average method:
[0172] I R =αI RGB +(1-α)I T
[0173] Finally, I R This is the pet food packaging bag image obtained after image registration processing.
[0174] 3. Image Segmentation Module Design
[0175] In order to achieve the segmentation of pet food packaging bag images, the present invention uses UNet as a segmentation model. It is used to extract the area of the pet food packaging bag and the printed pattern area therein. The traditional UNet model has problems such as blurred edges or inaccurate segmentation due to insufficient local features when dealing with complex packaging bag problems. The present invention introduces a dynamic context perception module and combines global context information of different scales to help improve the adaptability of the segmentation model to complex backgrounds. The overall process is as follows: Figure 4 shown.
[0176] Based on the image dataset I obtained after image registration R , using UNet as the segmentation model, from the input image dataset I R Middle split packaging bag area R p and the printed pattern area R s Specifically, UNet consists of two main parts: encoder and decoder. The encoder is responsible for extracting features and gradually reducing the resolution, while the decoder is responsible for restoring the spatial resolution of the image.
[0177] In the encoder, it consists of a series of convolutional layers and pooling layers, the main purpose of which is to gradually extract high-level image features while reducing the spatial resolution of the feature map.
[0178] In each layer, the input image is convolved, the convolution kernel size is set to 3×3, and nonlinearity is introduced through the activation function ReLU. R As the input of layer 0, it undergoes the convolution operation:
[0179]
[0180] Represents the output feature map after the convolution operation. are the parameters of the convolution kernel.
[0181] Next, the pooling operation is performed, which is used for downsampling. Dynamic context modeling is introduced to improve the adaptability of the segmentation model to complex backgrounds. The adaptive pooling operation generates multi-scale context features:
[0182]
[0183] In the decoder, the goal of the decoder part is to restore the low-resolution feature map to the size of the original image and combine the high-resolution features of the encoder part through skip connections.
[0184] Map the multi-scale context features back to the original resolution through weighted fusion. Upsample the feature map through deconvolution operation:
[0185]
[0186] is the upsampled feature map, the size is gradually restored, and l represents the lth layer. In the decoder, the jump connection concatenates the high-resolution features of the encoder with the upsampled results of the decoder to preserve the details:
[0187]
[0188] in, is the concatenated feature map, is a learnable weight parameter used to dynamically adjust the contribution of contexts of different scales.
[0189] Furthermore, the decoder convolves the concatenated feature maps and performs nonlinear transformation through ReLU activation:
[0190]
[0191] Finally, the decoder output F fin The number of channels is reduced to 1 through a 1×1 convolution layer, and the category probability of each pixel is mapped to the interval [0, 1] through the Sigmoid activation function, so that each pixel belongs to the bag area R p and the printed pattern area R s Probability
[0192]
[0193] in, is the output probability map, with a value of [0, 1], indicating the probability that each pixel belongs to its respective category. σ is the activation function, and k represents the kth image.
[0194] Next, by transforming the probability map Binarization to obtain the packaging bag area and printed pattern area Using threshold when Then the image (x, y) belongs to the packaging bag area, otherwise it belongs to the printing area.
[0195] Furthermore, UNet uses the cross entropy loss function to train the model, with the goal of minimizing the difference between the predicted results and the true labels. The formula of the cross entropy loss function is as follows:
[0196]
[0197] Among them, Y k (x, y) represents the true label, indicating whether the pixel (x, y) belongs to the target region. H and W are the height and width of the segmentation map, respectively. Through back propagation, the parameters of the model are updated to minimize the loss function
[0198] Through the above content, the segmented image I is finally obtained seg . As input to the pet food bag defect detection model.
[0199]
[0200] 4. Defect Detection Module Design
[0201] In the pet food packaging bag inspection task, the goal is to identify blurred printing, offset printing, and missing printing on the packaging bag. These defective areas need to be accurately identified for subsequent quality control.
[0202] According to the defect area image obtained by the segmentation module, the present invention adopts YOLOv5 as the target detection unit. seg As the model input image, it is used to identify the defect area in the image. Image I input to YOLOv5 segIt only contains the packaging bag and the printing area, so YOLOv5 can focus on the defect detection of these two areas. Since the feature extraction part of the network in YOLOv5 target detection mainly relies on the features of the single-modal input, it is difficult to make full use of the complementary information of multi-modal data. Therefore, the present invention adopts the feature mapping F = {F RGB , F T As a guide, a modality-guided dynamic feature enhancement unit is introduced to guide the RGB features through thermal imaging features, and dynamically adjust the feature weights of different areas, thereby improving the detection network's perception of defects.
[0203] The goal of the YOLOv5 target detection unit is to detect defective areas in the image, including printing blur, printing offset, and printing missing. seg After processing, a set of bounding boxes and class probabilities are output, indicating the location information and classification results of each defect in the image. Figure 4 shown.
[0204] The design introduces a modal-guided dynamic feature enhancement unit. Channel compression and global modeling are performed on thermal imaging features to generate guided weights:
[0205] w T =σ(FC(Gap(F T )))
[0206] Among them, Gap(·) represents global average pooling, which is used to extract global features, FC(·) represents the fully connected layer, which maps the global features to the weight space, and σ represents the Sigmoid activation function.
[0207] Use thermal imaging weights to adjust the channel importance of RGB features for dynamic feature enhancement:
[0208]
[0209] Furthermore, the enhanced RGB features are fused with the original thermal imaging features, and the convolution operation is used to further extract the local information in the fused features:
[0210]
[0211] Furthermore, using the YOLOV5 target detection unit and the enhanced feature F en To achieve defect detection of pet food packaging bags, the specific steps are as follows:
[0212] Will I seg The dataset is input into YOLOV5 for grid division and bounding box prediction; for the kth input image First, it is divided into s×s grids; then the bounding box prediction is performed for each grid unit. For the i-th network unit g i , YOLOV5 outputs the predicted bounding box B of the network unit i =(x i ,y i , w i ,h i ), (x i ,y i , w i ,h i ) is the position coordinate of the predicted bounding box; for the i-th bounding box, YOLOv5 will output a category probability C i , indicating the probability that the bounding box belongs to a certain type of defect; further, after the target detection is completed, the YOLOv5 model will generate a set of bounding boxes and category probabilities, and set non-maximum suppression (NMS) according to the detection results. That is, for multiple overlapping bounding boxes, non-maximum suppression will be based on the confidence threshold δ NMS Filter out the best bounding boxes and remove boxes with high overlap and low confidence.
[0213] Specifically, for the bounding box B i =(x i ,y i , w i ,h i ) and B j , calculate the intersection-over-union ratio of two bounding boxes:
[0214]
[0215] If the IoU of two bounding boxes is greater than the threshold δ NMS , the bounding boxes with higher confidence are retained and the bounding boxes with lower confidence are removed.
[0216] Adopt F en Calculate the confidence of the printing defects in the grid unit to correct the detection results of the YOLOV5 target detection unit; for the i-th network unit g i Defective confidence The feature pyramid network (FPN) is used to fuse features of different resolutions. en Input into the multi-scale feature layer of the pyramid network:
[0217]
[0218] in, is the original confidence value predicted by the model, is the normalized confidence in the interval [0, 1] obtained by the Sigmoid activation function;
[0219] The defect category probability output by YOLOv5 and the confidence output by the multi-scale feature layer of the pyramid network are combined and input into the softmax activation function to obtain the printing defect detection category class; for the i-th network unit g i , the specific calculation formula is as follows:
[0220]
[0221] The defects in the present invention are divided into three categories: printing blur (0), printing offset (1), and printing missing (2). i = 0 means that the area selected by the predicted bounding box in the i-th grid unit has a printing blur problem. i = 1 means that the area selected by the predicted bounding box in the i-th grid unit has a printing offset problem. i =2 means that there is a printing missing problem in the area selected by the predicted bounding box in the i-th grid unit.
[0222] 5. Model Training
[0223] The image registration module, image segmentation module and defect detection module are trained using the pet food packaging bag image data obtained through camera calibration and data collection. The process includes:
[0224] 1) Using the RGB image I obtained in S1 RGB and thermal imaging image I T , which together constitute the dataset I o , for the dataset I o Perform data enhancement, expand the dataset by image rotation, image scaling and noise addition, and merge it with the original image data to obtain the total image dataset I f .
[0225] 2) Using the total dataset I f Train the image registration module; first perform feature extraction, and extract the high-dimensional features of the image through the pre-trained convolutional neural network to obtain F RGB , F T Secondly, for F RGB , F T Perform feature point matching, find key points P, and generate M by matching point pairs. Optimize the matching point pairs M, remove the wrong matching point pairs, and obtain M F , the coordinates of the matching points (x, y) and (x′, y′) of the RGB and thermal imaging images are obtained through the homography transformation matrix. The corresponding image set is I R . Image registration is completed.
[0226] 3) Image dataset I obtained by image registration training R Train the image segmentation module. R As input to the UNet segmentation model for training, the UNet segmentation model consists of two parts, the encoder and the decoder. In the decoder part, the output feature map is obtained by introducing nonlinearity through convolution operations and activation functions. Then, a pooling operation is performed for downsampling. In order to improve the adaptability of the segmentation model to complex backgrounds, dynamic context modeling is introduced, and adaptive pooling operations are performed to generate multi-scale context features. right The upsampling is performed through weighted fusion deconvolution to realize the result splicing of the decoder part. The output F of the decoder fin The activation function maps the category of each pixel to the interval [0, 1] to obtain the probability of the category to which the pixel belongs. according to Determine it as the packaging bag area R p Or printing area R s The image dataset after image registration consists of I R Represented as I seg .
[0227] 4) Image dataset I obtained using the image segmentation module seg Train the defect detection module. Introduce the modality-guided dynamic feature enhancement module to train the feature map F of the RGB image obtained by the image registration module. T Perform channel compression and global modeling to generate guidance weights w T . Using the guided weight w T Dynamic feature enhancement of RGB image feature map F RGB ,get The two enhanced feature maps are fused, and the local information of the fused features is extracted using the convolution operation to obtain F en , and use it as the detection head input of the target detection model YOLOv5 for model training. en Input into the multi-scale feature layer of the pyramid network to obtain the original confidence value predicted by the model, and then train the activation function to obtain the confidence value of the target existence. and category scores
[0228] 6. Model deployment and application
[0229] The trained image registration module, image segmentation module and defect detection module are deployed on the server background to realize real-time detection of the printing quality of pet food packaging bags.
[0230] After the RGB image and thermal image of the pet food packaging bag are collected by the RGB camera and the thermal imaging camera respectively, the following steps are performed to complete the inspection and verification of the printing quality of the pet food packaging bag and the grading of the degree of verification:
[0231] Standardize and resize RGB and thermal images of pet food packaging bags to ensure that they meet the resolution and size requirements when input into the deployed model.
[0232] The pre-processed pet food packaging bag RGB image and thermal imaging image are input into the trained image registration module to obtain the registered image;
[0233] The registered image is input into a trained image segmentation module for image segmentation to obtain a pet food packaging bag area image and a printing area image;
[0234] The pet food packaging bag area image and the printing area image are input into the defect detection module to output the pet food packaging bag printing quality detection result.
[0235] Experimental evaluation and comparison: The algorithm of the present invention is compared with the traditional YOLO series target detection experiment, such as Figure 5 As shown. The algorithm proposed in the present invention performs best in the pet food packaging bag detection task, with an accuracy of 93%, a recall rate of 91%, and a comprehensive F1 score of 92%, which is significantly better than YOLOv3 (F1 score 82.5%) and YOLOv5 (F1 score 89%). Compared with the traditional YOLOv5, the multi-scale improvement enhances the detection ability of small targets and complex backgrounds, effectively improves the recall rate and accuracy, and is particularly suitable for the detection scenarios of printing defects and abnormal details. This shows that the method designed by the present invention is an effective solution for pet food packaging bag image registration printing quality detection.
[0236] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0237] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A pet food packaging bag image registration detection method, characterized in that: The following steps are involved: S1, collects RGB images and thermal imaging images of pet food packaging bags through RGB camera and thermal imaging camera respectively; S2, standardizes and preprocesses the RGB image and thermal image of the pet food packaging bag, and interpolates the thermal image to the same resolution as the RGB image; S3, inputting the pre-processed pet food packaging bag RGB image and thermal imaging image into the trained image registration model, and outputting the pet food packaging bag printing quality inspection result; The image registration model includes an image registration module, an image segmentation module and a defect detection module; The image registration module performs feature extraction, feature point matching, matching point pair optimization and fusion on the collected RGB image and thermal imaging image to generate a pixel-level fused image; The image segmentation module uses UNet as a segmentation model to segment the packaging bag area and the printed pattern area from the input pixel-level fusion image. The UNet consists of an encoder and a decoder. The encoder is responsible for extracting features and gradually reducing the resolution, while the decoder is responsible for restoring the spatial resolution of the image. The defect detection module performs detection based on the defect area image obtained after segmentation, and outputs the position information and classification result of each printing quality defect.
2. The pet food packaging bag image registration detection method according to claim 1, characterized in that: The process of constructing the dataset used to train the image registration model is as follows: Collect samples of pet food packaging bags from different batches, including normal samples and defective samples; the types of defects include missing printing, offset printing, and blurred printing; The images collected by the RGB camera and the thermal imaging camera have the same viewing angle and uniform brightness. The acquisition frequency and the acquisition time of each sample are designed to ensure that the complete area of the packaging bag is covered; The printing defects in the RGB image and the thermal imaging image are annotated, and the defects in the printed part of the pet food packaging in the image are annotated with a rectangular box, and the rectangular bounding box is B = {x min ,y min , x max ,y max },(x min ,y min ) and (x max ,y max ) represent the coordinates of the upper left corner and the lower right corner of the bounding box respectively; defects are divided into three types: printing blur, printing offset and printing missing, which are represented by codes 0, 1 and 2 respectively.
3. The pet food packaging bag image registration detection method according to claim 1, characterized in that: The feature extraction in the image registration module is to perform multi-layer feature extraction, where each layer uses a pre-trained convolutional neural network to extract high-dimensional features of RGB and thermal imaging images: F RGB =f CNN (I RGB ) F T =f CNN (I T1 ) Among them, f CNN represents the convolutional neural network, F RGB and F T are the feature maps of RGB images and thermal images respectively; the dimension of the output feature map is F∈R H′×W′×C , where H′ and W′ represent the height and width of the feature map, which is the downsampling of the input image resolution; C represents the number of feature channels, which represents the extracted feature dimension; The pre-trained convolutional neural network includes a convolution layer, a pooling layer, and a normalization layer; The convolutional layer is used to extract local features: Among them, F l (x, y, c) represents the value of channel c at position (x, y) in the output feature map of the lth layer, i.e., the output feature map of the lth layer; W l (i, j, c) represents the convolution kernel weight, and the kernel size is (2k+1)×(2k+1); b i is the bias term, k is the convolution kernel radius; The pooling layer uses the maximum pooling operation to downsample and reduce the spatial resolution of the feature map: Where R represents the pooling window range, x′, y′ represent the coordinates after pooling; Finally, batch normalization and nonlinear activation function ReLU are used to enhance the feature expression ability and obtain a multi-layer feature map F RGB and F T .
4. The pet food packaging bag image registration detection method according to claim 3, characterized in that: The feature point matching in the image registration module is to achieve cross-modal key point alignment through feature point detection and matching, and establish the correspondence between multi-modal images: First, key point extraction is performed; key points P are detected in each feature map: P={p i |i=1,2,...,N} Wherein, each key point p includes two-dimensional coordinates (x, y) and feature descriptor d, and N is the number of key points detected in each feature map; Next, a feature descriptor d is generated for each key point, extracted from its local feature map: d(p)=desc(F,p) Where desc(F, p) is expressed as RGB , F T } extract the local feature descriptor with p as the center; Secondly, generate matching point pairs; use the similarity measure between feature descriptors to match key points in different feature maps: Wherein, sim(d1, d2) represents a similarity measurement function, the present invention adopts Euclidean distance calculation, τ is a similarity threshold, is the i-th pixel in the RGB image, is the jth pixel in the thermal imaging image.
5. The pet food packaging bag image registration detection method according to claim 4, characterized in that: The matching point pair optimization and fusion in the image registration module first use the random sampling consensus algorithm RANSAC to eliminate the wrong matching point pairs and retain the internal point set. f The set of matched point pairs after optimization is: M f =RANSAC(M) Then, the homography transformation matrix is estimated based on the optimized matching point pair M f , estimate the homography transformation matrix H from thermal imaging to RGB image: Among them, h ** Respectively represent the effects of transformation and perspective transformation in the x direction and y direction; the homography transformation relationship is: Where (x, y) and (x′, y′) are the coordinates of the matching points in the thermal image and RGB image respectively; Finally, pixel-level image fusion is performed. During the fusion process, the pixel values of multiple images are merged into a final image. Set the weight coefficient α, according to the weighted average method: I R =αI RGB +(1-α)I T Finally, I R is the registered image obtained after image registration processing.
6. The pet food packaging bag image registration detection method according to claim 1, characterized in that: The image segmentation module adopts the UNet model, in which the encoder consists of a series of convolutional layers and pooling layers; At each layer, the input registered image is convolved, and nonlinearity is introduced through the activation function ReLU. R As the input of layer 0, it undergoes the convolution operation: Represents the output feature map after the convolution operation; is the parameter of the convolution kernel; Next, the pooling operation is performed. The pooling operation is used for downsampling, dynamic context modeling is introduced, and the adaptive pooling operation generates multi-scale context features: The decoder aims to restore the low-resolution feature map to the size of the original image and combine the high-resolution features of the encoder part through skip connections; The multi-scale context feature maps are mapped back to the original resolution, and the feature maps are upsampled by deconvolution operations through weighted fusion: is the upsampled feature map, the size is gradually restored, l represents the lth layer; in the decoder, the jump connection splices the high-resolution features of the encoder with the upsampled results of the decoder to retain the details: in, is the concatenated feature map, is a learnable weight parameter used to dynamically adjust the contribution of contexts of different scales; The decoder convolves the concatenated feature map and performs nonlinear transformation through ReLU activation: Finally, the decoder output F fin The number of channels is reduced to 1 through a 1×1 convolution layer, and the category probability of each pixel is mapped to the interval [0, 1] through the Sigmoid activation function, so that each pixel belongs to the bag area R p and the printed pattern area R s Probability in, is the output probability map, with a value of [0, 1], indicating the probability that each pixel belongs to its respective category, σ is the activation function, and k represents the kth image; Next, by transforming the probability map Binarization to obtain the packaging bag area and printed pattern area Using threshold when Then the image (x, y) belongs to the packaging bag area, otherwise, it belongs to the printing area, and finally the segmented image I is obtained. seg .
7. The pet food packaging bag image registration detection method according to claim 6, characterized in that: The UNet model uses the cross entropy loss function to train the model, the goal is to minimize the difference between the predicted results and the true labels. The formula of the cross entropy loss function is as follows: Among them, Y k (x, y) represents the true label, indicating whether the pixel (x, y) belongs to the target area. H and W are the height and width of the segmentation map, respectively. Through back propagation, the parameters of the model are updated to minimize the loss function.
8. The pet food packaging bag image registration detection method according to claim 1, characterized in that: The defect detection module includes two parts: a dynamic feature enhancement unit that introduces modality guidance and a YOLOv5 target detection unit; The dynamic feature enhancement unit performs channel compression and global modeling on the thermal imaging features to generate guidance weights: W T =σ(FC(Gap(F T ))) Among them, Gap(·) represents global average pooling, which is used to extract global features, FC(·) represents the fully connected layer, which maps the global features to the weight space, and σ represents the Sigmoid activation function; Use thermal imaging weights to adjust the channel importance of RGB features for dynamic feature enhancement: The enhanced RGB features are fused with the original thermal imaging features, and the convolution operation is used to further extract the local information in the fused features: Using the YOLOv5 target detection unit and the enhanced feature F en Enables defect detection in pet food packaging bags.
9. The pet food packaging bag image registration detection method according to claim 8, characterized in that: Using the YOLOv5 target detection unit and the enhanced feature F en The specific process of realizing defect detection of pet food packaging bags is as follows: The segmented image I seg Input into the YOLOV5 target detection unit for grid division and bounding box prediction; for the kth input image First, it is divided into s×s grids; then the bounding box prediction is performed for each grid unit. For the i-th network unit gi, the YOLOV5 target detection unit outputs the predicted bounding box B of the network unit i =(x i ,y i , w i ,h i ), (x i ,y i , w i ,h i ) is the position coordinate of the predicted bounding box; for the i-th bounding box, YOLOv5 will output a category probability C i , represents the probability that the bounding box belongs to a certain type of defect; After the target detection is completed, a set of bounding boxes and category probabilities are generated. According to the detection results, non-maximum suppression NMS is set. For multiple overlapping bounding boxes, NMS will be based on the confidence threshold δ NMS Filter out the best bounding boxes and remove boxes with high overlap and low confidence; For the bounding box B i =(x i ,y i , w i ,h i ) and B j , calculate the intersection of two bounding boxes: If the IoU of two bounding boxes is greater than the threshold δ NMS , then the bounding box with higher confidence is retained and the bounding box with lower confidence is removed; Adopt F en Calculate the confidence of the printing defects in the grid unit to correct the detection result of YOLOV5; for the i-th network unit g i Defective confidence Use the feature pyramid network FPN to fuse features of different resolutions. en Input into the multi-scale feature layer of the pyramid network: in, is the original confidence value predicted by the model, is the normalized confidence in the interval [0, 1] obtained by the Sigmoid activation function; The defect category probability output by YOLOv5 and the confidence output by the multi-scale feature layer of the pyramid network are combined and input into the softmax activation function to obtain the printing defect detection category class; for the i-th network unit g i , the specific calculation formula is as follows: Defects are divided into three categories: printing blur 0, printing offset 1, and printing missing 2. If class i = 0 means that the area selected by the predicted bounding box in the i-th grid unit has a printing blur problem. i = 1 means that the area selected by the predicted bounding box in the i-th grid unit has a printing offset problem. i =2 means that there is a printing missing problem in the area selected by the predicted bounding box in the i-th grid unit.
Citation Information
Patent Citations
Electrical equipment appearance abnormity detection method based on image comparison
CN104809732A
Geosynchronization of an aerial image using localizing multiple features
WO2024042508A1
Cited By
Double-outlet square bag zippered bag making detection method and system based on image detection
CN120765551A
Intelligent active package optimization method and system based on machine learning
CN120893631A
Non-setting adhesive product detection device and method
CN121437391A
Image processing-based hose coupler assembly deformation detection method
CN121685400A
Intelligent inspection method and system based on AI image recognition
CN121686252A