Bimodal industrial packaging defect detection method

Through the dual-mode industrial packaging defect detection method, combined with deep learning and optical character recognition technology, the existing tag defect detection methods are solved, and the problems of inefficient, poor adaptability and insufficient accuracy are realized, and the real comparison and defect detection of real tag images and electronic design drawings are realized, which improves detection efficiency and accuracy, and is suitable for high-speed production line environments.

CN120219385APending Publication Date: 2025-06-27SOUTHWEAT UNIV OF SCI & TECH

Patent Information

Application Number
CN202510692220.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing label defect detection methods have problems such as inefficiency, poor adaptability and insufficient accuracy, especially when packaging labels of different sizes and shapes, high-speed production lines and label content are complex.

Method used

The dual-modal industrial packaging defect detection method is adopted, combined with image acquisition, deep learning, optical character recognition (OCR) and template matching technology, text images and label area images are separated through the YOLOv11s model, and text defects are identified by enhanced real-time scene text detection network and scene text recognition model, and label area image defects are detected based on grayscale matching method.

Benefits of technology

It realizes rapid comparison of real-life images of tags and electronic design drawings of tags, improves the efficiency and accuracy of tag defect detection, enhances the adaptability and robustness of algorithms, and can run in real time in a high-speed production line environment, meeting the high-precision and high-efficiency needs of modern industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219385A_ABST
    Figure CN120219385A_ABST
Patent Text Reader

Abstract

The invention discloses a bimodal industrial package defect detection method, and belongs to the technical field of industrial package label quality detection, and the method comprises the steps: collecting a label electronic design drawing and a label object image on an industrial package; respectively separating the label physical image and the label electronic design drawing into a text image and a label area image; identifying a text region in the text image by adopting an enhanced real-time scene text detection network; recognizing a text region by adopting a scene text recognition model, and detecting text defects; detecting label area image defects based on a gray scale matching method; and fusing the text image and the label area image to the original position, and marking the positions of the image defect and the text defect. According to the invention, the efficiency and precision of label defect detection are improved, the adaptability and robustness of the algorithm are significantly enhanced through the adaptive image acquisition module and the bimodal detection method, the requirements of high precision and high efficiency of modern industrial production are met, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of quality inspection of industrial packaging labels, and particularly relates to a dual-modal industrial packaging defect detection method. Background Art

[0002] With the development of industrial automation and intelligence, label quality management on the production line has become an important link in modern production. As the identifier of packaging, labels carry important product information, such as text, graphics, barcodes, two-dimensional codes, etc. Incorrect or defective labels not only affect product quality and brand image, but may also lead to unqualified products flowing into the market, causing serious economic losses.

[0003] Currently, most traditional label defect detection methods rely on manual inspection, which is time-consuming and laborious, and prone to missed detections and misjudgments. With the progress of artificial intelligence technology, especially the development of computer vision and deep learning, automated defect detection based on image processing and pattern recognition has gradually become the mainstream. However, the existing label defect detection methods have the following problems:

[0004] (1) Limitations of fixed template matching: Traditional vision detection systems often rely on preset templates for matching, but for packaging labels of different sizes and shapes, template matching cannot adapt to changes, resulting in missed detections and false alarms.

[0005] (2) Limitations of single-scale detection networks: Single-scale detection methods are prone to missed detections and inaccurate results when facing different shooting distances, especially when the text of small-sized labels is blurred.

[0006] (3) Processing delay in high-speed production lines: In the scenario of high-speed production lines, the processing speed of existing detection systems often becomes a bottleneck, resulting in real-time detection and feedback being unable to meet production requirements.

[0007] (4) Complex label content: Most labels contain various types of targets to be recognized, such as text, images, special characters, etc., making the detection difficult. Summary of the Invention

[0008] The purpose of the present invention is to provide a dual-modal industrial packaging defect detection method for the above-mentioned deficiencies in the prior art, so as to solve the problems of low efficiency, poor adaptability, and insufficient accuracy in traditional label defect detection methods.

[0009] To achieve the above object, the technical solution adopted by the present invention is:

[0010] A dual-modal industrial packaging defect detection method, which includes the following steps:

[0011] S1. Collect the electronic design drawings of the labels and the physical images of the labels on the industrial packaging;

[0012] S2. Use the YOLOv11s model to separately separate the label physical image and the label electronic design drawing into a text image and a label area image respectively;

[0013] S3. Use the enhanced real-time scene text detection network to identify the text area in the text image;

[0014] S4. Use the scene text recognition model to recognize the text area and detect text defects;

[0015] S5. Detect label area image defects based on the gray-scale matching method;

[0016] S6. Merge the two text images and the two label area images to the original positions of the label physical image and the label electronic design drawing, and mark the positions of the image defects and text defects.

[0017] Furthermore, S2 specifically includes:

[0018] Use the trained YOLOv11s model to recognize the image areas in the label physical image and the label electronic design drawing, and then separately separate the label physical image and the label electronic design drawing into a text image and a label area image respectively. At the same time, record the position information of the text image and the label area image.

[0019] Furthermore, S3 includes the following sub-steps:

[0020] S31. Input the text image into the ResNet101 network in the enhanced real-time scene text detection network to output 4 feature maps of different scales;

[0021] S32. Input the 4 feature maps of different scales into the feature pyramid module in the enhanced real-time scene text detection network to output 4 enhanced multi-scale feature maps;

[0022] S33. Input the 4 enhanced multi-scale feature maps into the adaptive scale fusion module in the enhanced real-time scene text detection network to output a fused feature;

[0023] S34. Perform a differentiable binarization operation on the fused feature to generate a binarized mask of the text area.

[0024] Furthermore, in S31, output 4 feature maps of different scales, including:

[0025]

[0026] where is the i-th feature map of different sizes, including the feature maps F1, F2, F3, and F4 of different scales; Denote the ResNet101 network; I is the text image.

[0027] Furthermore, in S32, four enhanced multi-scale feature maps are output, including:

[0028] In the feature pyramid module, layer-by-layer sampling is performed sequentially from the feature maps F4, F3, F2, and F1 and fused with adjacent low-resolution feature maps:

[0029]

[0030]

[0031]

[0032] In the formula, , , are multi-scale feature maps; is the upsampling operation;

[0033] The feature maps F4, multi-scale feature maps , , are upsampled to the same size through upsampling to obtain the enhanced multi-scale feature maps , , and .

[0034] Furthermore, S33 includes the following sub-steps:

[0035] S331. Connect the input multi-scale feature maps , , and , and divide them into two branches. Perform Conv operations on the two branches respectively to obtain the intermediate feature S;

[0036] S332. In one branch, apply the spatial attention mechanism to the intermediate feature S to obtain the spatial attention weight :

[0037]

[0038] In the formula, is the spatial attention mechanism;

[0039] In the other branch, apply the channel attention to the intermediate feature S to directly obtain the channel attention weight :

[0040]

[0041] In the formula, is the channel attention mechanism;

[0042] S333. Divide the spatial attention weight and the channel attention weight along the channel dimension into 4 parts, and multiply them with the corresponding feature weights to obtain the fused feature:

[0043]

[0044] In the formula, F is the fused feature; is the spatial attention weight of the i-th part; is the channel attention weight of the i-th part; is the i-th enhanced multi-scale feature map.

[0045] Furthermore, in S34, perform a differentiable binarization operation on the fused feature F to generate a binarized mask for the text region, including:

[0046]

[0047] In the formula, k is a hyperparameter that controls the slope of the binarization curve; t is the binarization threshold; B is the binarized mask, representing the text region; e is the natural base.

[0048] Furthermore, S4 specifically includes:

[0049] Input the text image corresponding to the text region into the scene text recognition model, convert the text image into character components through hierarchical overlapping patch embedding, and perform three-stage feature extraction on the character components at different scales;

[0050] Among them, each stage uses a hybrid block for feature extraction:

[0051]

[0052] In the formula, is the output feature map of the i-th hybrid block, ; is the input feature map of the i-th hybrid block; is the hybrid block operation;

[0053] Perform a merging operation on the output feature map :

[0054]

[0055] In the formula, C is the merged feature map; is the merging operation;

[0056] The merged feature map C is used to recognize each character through a linear prediction layer, and the recognized characters are compared with the standard text content to detect text defects.

[0057] Furthermore, S5 specifically includes:

[0058] Calculate the structural similarity index SSIM between the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing, and then detect the label area image defect:

[0059]

[0060] In the formula, and are the average gray values of the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing respectively; and are the gray variances of the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing respectively; is the gray covariance of the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing; and are constants.

[0061] The dual-modal industrial packaging defect detection method provided by the present invention has the following beneficial effects:

[0062] 1. The present invention combines image acquisition, deep learning, optical character recognition (OCR) and template matching technologies to realize rapid comparison between the label physical image and the label electronic design drawing, so as to quickly detect text and image defects on the label.

[0063] 2. The present invention not only improves the efficiency and accuracy of label defect detection, but also significantly enhances the adaptability and robustness of the algorithm through the adaptive image acquisition module and the dual-modal detection method. And the experimental results show that the present invention can run in real time in a high-speed production line environment, meet the high-precision and high-efficiency requirements of modern industrial production, and has broad application prospects. Description of the Drawings

[0064] Figure 1 is the flowchart of the dual-modal industrial packaging defect detection method of the present invention.

[0065] Figure 2 is the image acquisition schematic diagram of the 9-eye collaborative adaptive image acquisition workbench of the present invention.

[0066] Figure 3Schematic diagram of non - maximum suppression for Canny edge detection of the present invention.

[0067] Figure 4 Schematic diagram of the text image and label area image of the present invention.

[0068] Figure 5 Structural diagram of the EDBNet++ network of the present invention.

[0069] Figure 6 Structural diagram of the SVTR network of the present invention.

[0070] Figure 7 Schematic diagram of the template matching result of the present invention. Detailed implementation manners

[0071] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0072] Example 1

[0073] This example provides a dual - mode industrial packaging defect detection method, which can effectively improve the automation level of label quality management, reduce manual intervention, lower production costs, and improve product quality and brand image. Refer to Figure 1 , and it specifically includes the following contents:

[0074] S1. Collect the physical image of the label on the industrial packaging and the electronic design drawing of the label;

[0075] Specifically, in this example, to meet the acquisition requirements of labels of different sizes, a 9 - camera collaborative adaptive image acquisition workbench is designed, as shown in Figure 2 ; this workbench can automatically adjust the acquisition parameters according to the size and position of the label to ensure the clarity and integrity of the image. The specific implementation method is as follows:

[0076] Camera layout and control: Nine industrial - grade high - definition cameras are arranged on a fixed workbench. The positions and angles of the cameras are precisely calculated to ensure coverage of the entire physical label area. The cameras are arranged in a matrix layout. The central camera is mainly responsible for collecting the core area of the label, while the peripheral cameras are responsible for capturing the edge and corner areas of the label. Through this layout, regardless of the size of the physical label, all label information can be accurately captured by the cameras within a certain range.

[0077] Automatic label size recognition: The workbench is equipped with a size recognition system that automatically recognizes the size of the label through image processing algorithms (for example, using the preliminary image of the label or fiducial marks for recognition). When the label is placed on the workbench, the size recognition system detects the contour, size, and position of the label through vision algorithms to determine the specific size of the label.

[0078] Adaptive camera angle adjustment: According to the size information of the label, the focal length and exposure settings of the camera are automatically adjusted to ensure the clarity and accuracy of the image. The focal length and exposure of the camera are automatically adjusted to ensure that the entire physical image of the label can be clearly presented simultaneously, reducing unnecessary background interference.

[0079] Image stitching and optimization: For large-sized labels, the images from multiple cameras are merged into a complete label image through image stitching algorithms. The stitching algorithm is accurately aligned based on the identifying features in the image (such as the edges of the label, alignment points, etc.) to ensure seamless connection of the stitched image. For small-sized labels, the image is captured by the central camera, and the image quality is automatically optimized to ensure that there is no excessive background noise in the image.

[0080] Dynamic adjustment of image acquisition parameters: To adapt to the acquisition of different-sized labels in different environments, an adaptive parameter adjustment module is built-in. When the label size or environmental light changes, the adaptive parameter adjustment module adjusts the exposure, white balance, contrast, and other parameters of each camera in real time to ensure that the acquired images can maintain high quality in strong light, low light, or shadow areas.

[0081] The 9-camera collaborative adaptive image acquisition workbench in this embodiment can flexibly handle the image acquisition requirements of different-sized labels, avoiding the complexity of using different acquisition devices for small and large labels. At the same time, it also improves the efficiency and accuracy in the acquisition process to obtain a physical label image that meets the accuracy and size requirements. For the physical label image, the 9-camera collaborative adaptive image acquisition workbench is used for image acquisition to obtain the image data corresponding to the physical label; at the same time, for the electronic design drawing of the label, it can be directly converted into a picture format, so that the physical label image data and the electronic design drawing image data of the label are obtained.

[0082] S2. Use the YOLOv11s model to separate the physical label image into a text image and a label area image, and at the same time separate the electronic design drawing of the label into a text image and a label area image;

[0083] Specifically, in this embodiment, the collected physical images of the tags and the electronic design drawings of the tags are first subjected to Canny edge detection to determine their relative sizes, facilitating subsequent operations. The Canny algorithm is a standard algorithm widely used in edge detection, and its goal is to find an optimal edge detection solution or locate the positions with the strongest gray intensity changes in an image.

[0084] During the edge detection process, first, the image is subjected to Gaussian filtering to remove noise and smoothed using a Gaussian filter. The convolution kernel expression of the Gaussian filter is as follows:

[0085]

[0086] In the formula, is the standard deviation of the Gaussian function, is the weight of the Gaussian kernel, and x and y are the abscissa and ordinate of the convolution kernel respectively.

[0087] Calculate the gradients of the image in the horizontal and vertical directions, and obtain the gradient intensity and direction through the Sobel operator; the calculation formula of the Sobel operator is:

[0088]

[0089] In the formula, is the gradient in the x direction; is the gradient in the y direction; is the input image;

[0090] The gradient intensity G and direction θ can be expressed as:

[0091]

[0092] Next, referring to Figure 3 , non-maximum suppression is performed, scanning the entire gradient intensity image and removing non-local maxima in the gradient direction, retaining local maxima, as shown in Figure 3 . Finally, double-threshold detection is performed, applying two thresholds to mark strong edges and weak edges respectively. Strong edges are pixels with gradient intensity greater than the high threshold, and weak edges are pixels with gradient intensity between the low threshold and the high threshold. By connecting weak edges to strong edges, complete edges are formed.

[0093] Considering the complexity of the label content, the label content is divided into two major categories: text and images, and a dual-modal detection method is adopted to conduct a comprehensive detection of the labels. For the text and image parts of the label physical image and the label electronic design drawing, different deep learning algorithms are applied for processing. Considering that some images may contain text information, if only text detection is used for differentiation, some images may be misdetected as text content. Therefore, this embodiment adopts an innovative separation strategy:

[0094] Reference Figure 4 , the YOLOv11s model is used to train the images in the label, and then the YOLOv11s model after completing the image training is used to identify the image regions in the label physical image and the label electronic design drawing. Furthermore, the label physical image is separated into a text image and a label region image. At the same time, the label electronic design drawing is also separated into a text image and a label region image, and the position information of the text image and the label region image is recorded, laying a foundation for subsequent text-image fusion.

[0095] S3. Use an enhanced real-time scene text detection network to identify the text regions in the text image, which specifically includes the following sub-steps:

[0096] S31. In this embodiment, an enhanced real-time scene text detection network (Enhanced Real-Time SceneText Detection with Differentiable Binarization and Adaptive Scale Fusion, EDBNet++) is constructed, and its network structure is as Figure 5 shown;

[0097] The separated text image is input into the ResNet101 network in the enhanced real-time scene text detection network (EnhancedReal-Time Scene Text Detection with Differentiable Binarization and AdaptiveScale Fusion, EDBNet++). By using its deep structure and residual connection mechanism, the accuracy of feature extraction can be effectively improved, and four feature maps of different scales are output;

[0098] Output four feature maps of different scales, including:

[0099]

[0100] where is the i-th feature map of different sizes, including feature maps of different scales 、 , , , where H, W, and C represent the height, width, and number of channels of the image respectively, represent the number of channels in different stages respectively; represents the ResNet101 network; I is the text image.

[0101] S32. Input the feature maps of 4 different scales into the Feature Pyramid Networks (FPN) in the enhanced real-time scene text detection network. This module further enhances the multi-scale expression ability of the feature maps through a bottom-up path and outputs 4 enhanced multi-scale feature maps;

[0102] In the feature pyramid module, perform layer-by-layer sampling from the feature maps F4, F3, F2, F1 with the highest resolution and fuse them with the adjacent low-resolution feature maps. The formula for the bottom-up path is:

[0103]

[0104]

[0105]

[0106] In the formula, , , are multi-scale feature maps; is the upsampling operation;

[0107] Convert the feature map F4 and the multi-scale feature maps , , to the same size through upsampling to obtain the enhanced multi-scale feature maps , , and .

[0108] S33. Input the 4 enhanced multi-scale feature maps into the Adaptive Scale Fusion (ASF) module in the enhanced real-time scene text detection network to output the fused features; The ASF module combines the Spatial Attention Mechanism (SAM) and the Channel Attention Mechanism (CAM). The spatial attention mechanism enables the network to focus on the key regions in the image, while the channel attention mechanism allows the network to learn the importance of different feature channels and assigns weights to each feature position. Assume that the input feature map consists of N feature maps It consists of N = 4. During specific processing, it includes the following sub-steps:

[0109] S331. Concatenate the input multi-scale feature maps , , and , and divide them into two branches. Perform Conv operations on the two branches respectively to obtain the intermediate feature S;

[0110]

[0111] In the formula, is the convolution operation, is the concatenation operation;

[0112] S332. In one branch, apply the spatial attention mechanism to the intermediate feature S to obtain the spatial attention weight :

[0113]

[0114] In the formula, is the spatial attention mechanism;

[0115] In the other branch, directly apply the channel attention to the intermediate feature S to obtain the channel attention weight :

[0116]

[0117] In the formula, is the channel attention mechanism;

[0118] S333. Divide the spatial attention weight and the channel attention weight into 4 parts along the channel dimension, and multiply them with the corresponding features weighted to obtain the fused feature:

[0119]

[0120] In the formula, F is the fused feature; is the spatial attention weight of the i-th part; is the channel attention weight of the i-th part; is the i-th enhanced multi-scale feature map.

[0121] S34. Perform a differentiable binarization operation on the fused feature to generate a binarized mask for the text region.

[0122]

[0123] Wherein, k is a hyperparameter for controlling the slope of the binarization curve; t is the binarization threshold; B is the binarization mask, representing the text region; e is the natural base number.

[0124] Through the method of Embodiment S3, the EDBNet++ network can not only learn the global text features, but also accurately locate the local key information, thereby achieving higher accuracy and robustness in the text detection task.

[0125] S4. Use a scene text recognition model to recognize the text region and detect text defects;

[0126] In S3, the text regions in the image have been accurately framed and their coordinate information has been obtained. The text images corresponding to the text regions are input into a scene text recognition model (Scene Text Recognition with a Single Visual Model, SVTR), refer to Figure 6 , and the text images are converted into character components through hierarchical overlapping patch embedding, and three-stage feature extraction is performed on the character components at different scales;

[0127] Specifically, when a text image of size (such as "Dual-Modal Industrial Packaging Detection System") is input into the network, first, the text image is converted into character components P through hierarchical overlapping patch embedding:

[0128]

[0129] Wherein, is hierarchical overlapping patch embedding;

[0130] Each character component is associated with a small part of the text characters in the text image, which enables the SVTR model to capture the local features of the characters. This process is achieved by splitting the image and mapping each block to an independent feature vector, and the SVTR model can handle the complex details and different text structures in the image.

[0131] Then, the SVTR model performs three-stage (Stage 1, 2, 3) feature extraction at different scales. These feature extractions are completed through a series of mixing blocks, and each mixing block consists of a local mixing and a global mixing module. The local mixing module focuses on capturing the fine-grained features inside the characters, such as the shape and strokes of the characters, etc., while the global mixing module is responsible for modeling the dependencies between the characters, such as the arrangement order of the characters and the overall structure of the text. Let the input of each mixing block be the feature map , the output passing through each mixing block is as follows:

[0132]

[0133] In the formula, is the output feature map of the i-th mixing block, ; is the input feature map of the i-th mixing block; is the mixing block operation, which captures the fine-grained features inside the characters and the global dependencies between characters; in this way, the SVTR model can extract rich features at multiple scales to adapt to texts of different sizes and complexities.

[0134] After the mixing block at each stage, the feature map is downsampled through a merging operation, reducing the height while keeping the width unchanged. Through this process, the SVTR model can fuse information at different scales. Suppose the input feature map at the i-th stage is , and the merged feature map is C:

[0135]

[0136] In the formula, C is the merged feature map; is the merging operation;

[0137] The merged feature map C is passed through a linear prediction layer to recognize each character. This linear prediction layer generates a probability distribution for each character based on the previously extracted features. By analyzing these probability distributions, the most likely category of each character can be determined, thus completing the entire text recognition task and obtaining an accurate recognition result (such as "bimodal industrial packaging detection system"), comparing the recognized characters with the standard text content, and then detecting text defects.

[0138] For each text box detected in the image, the SVTR model repeats this loop operation to obtain the recognition results of all texts. By comparing these recognition results with the standard text content, potential defects in the image can be detected. Using the coordinate information of each text box, the positions of these defects can be accurately marked on the image. Through this systematic method, not only can efficient and accurate text recognition be achieved, but also the defects that may occur during the recognition process can be accurately located and analyzed, which provides a comprehensive and effective solution for the text recognition task.

[0139] S5. Detect label area image defects based on the gray-scale matching method;

[0140] The label area images separated from the physical label image and the label electronic design drawing by the YOLOv11 model are rectangular. It is necessary to perform edge detection on the separated images again. Given the relative sizes of the physical label image and the label electronic design drawing, image matching can be used to detect whether there are defects in the label area images. As Figure 7 shown, the standard template image in the label electronic design drawing is shown on the right side of the screen. The area marked by the red frame in this template is the feature. Through the template matching algorithm, the corresponding actual area can be accurately located in the physical label image, realizing the accurate comparison between the design drawing and the physical image.

[0141] Calculate the structural similarity index SSIM between the label area image x1 in the physical label image and the label area image y1 in the label electronic design drawing, and then detect the defects in the label area image:

[0142]

[0143] In the formula, and are the average gray values of the label area image x1 in the physical label image and the label area image y1 in the label electronic design drawing respectively; and are the gray variances of the label area image x1 in the physical label image and the label area image y1 in the label electronic design drawing respectively; is the gray covariance of the label area image x1 in the physical label image and the label area image y1 in the label electronic design drawing; and are constants.

[0144] By comparing the structural similarity index SSIM value, the difference between the physical label image and the label electronic design drawing can be judged, and then the defects in the image can be identified.

[0145] S6. Fusion the text image and the label area image to the original position, and mark the positions of image defects and text defects.

[0146] By calling the pre-stored coordinate information of the label area image and the text image, the separated text image and label area image are intelligently recombined to accurately restore them to the original position. Specifically, the corresponding text image and label area image are fused to the original position of the label electronic design drawing; at the same time, the corresponding text image and label area image are fused to the original position of the physical label image; and the positions of the defects are marked in their respective images.

[0147] In the physical image of the label, the defective area is marked with a prominent red rectangular frame to precisely indicate the defect location. Meanwhile, in the electronic design drawing of the label, a green rectangular frame is used to synchronously mark the corresponding defect location, facilitating cross-modal comparative analysis for users. This dual marking mechanism not only significantly improves the accuracy and reliability of defect identification but also optimizes the user interaction experience, enabling users to quickly locate the problem area and take corresponding measures, effectively enhancing work efficiency.

[0148] Although the specific embodiments of the invention have been described in detail in conjunction with the accompanying drawings, it should not be construed as a limitation on the protection scope of this patent. Within the scope described in the claims, various modifications and deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of this patent.

Claims

1. A dual-modal industrial packaging defect detection method, characterized in that It includes the following steps: S1. Collect the label electronic design drawings and the physical label images on industrial packages; S2. Use the YOLOv11s model to separately separate the physical label image and the label electronic design drawing into a text image and a label area image; S3. Use an enhanced real-time scene text detection network to identify the text area in the text image; S4. Use a scene text recognition model to recognize the text area and detect text defects; S5. Detect label area image defects based on the gray-scale matching method; S6. Fuse the two text images and the two label area images to the original positions of the physical label image and the label electronic design drawing, and mark the positions of image defects and text defects.

2. The dual-modal industrial packaging defect detection method according to claim 1, characterized in that, The specific content of S2 includes: Use the trained YOLOv11s model to identify the image areas in the physical label image and the label electronic design drawing, and then separately separate the physical label image and the label electronic design drawing into a text image and a label area image. At the same time, record the position information of the text image and the label area image.

3. The dual-modal industrial packaging defect detection method according to claim 1, characterized in that S3 includes the following sub-steps: S31. Input the text image into the ResNet101 network in the enhanced real-time scene text detection network to output 4 feature maps of different scales; S32. Input the 4 feature maps of different scales into the feature pyramid module in the enhanced real-time scene text detection network to output 4 enhanced multi-scale feature maps; S33. Input the 4 enhanced multi-scale feature maps into the adaptive scale fusion module in the enhanced real-time scene text detection network to output a fused feature; S34. Perform a differentiable binarization operation on the fused feature to generate a binarization mask of the text area.

4. The dual-modal industrial packaging defect detection method according to claim 3, wherein In S31, outputting 4 feature maps of different scales includes: In the formula, is the feature map of the i-th different size, including feature maps F1, F2, F3, and F4 of different scales; represents the ResNet101 network; I is the text image.

5. The dual-modal industrial packaging defect detection method according to claim 4, wherein In S32, outputting 4 enhanced multi-scale feature maps includes: In the feature pyramid module, perform layer-by-layer sampling from the feature maps F4, F3, F2, and F1 in sequence and fuse them with the adjacent low-resolution feature maps: In the formula, , , are multi-scale feature maps; is an upsampling operation; The feature map F4 and the multi-scale feature maps , , are upsampled to the same size to obtain the enhanced multi-scale feature maps , , and .

6. The dual-modal industrial packaging defect detection method according to claim 5, wherein, S33 includes the following sub-steps: S331. Concatenate the input multi-scale feature maps , , and , divide them into two branches, and perform Conv operations on the two branches respectively to obtain the intermediate feature S; S332. In a branch, the spatial attention mechanism is applied to the intermediate feature S to obtain the spatial attention weight : In the formula, is the spatial attention mechanism; In another branch, the channel attention is directly applied to the intermediate feature S to obtain the channel attention weight : In the formula, is the channel attention mechanism; S333. Divide the spatial attention weight and the channel attention weight into 4 parts along the channel dimension, and multiply them with the corresponding feature weights to obtain the fused feature: Where F is the fused feature; is the spatial attention weight of the i-th part; is the channel attention weight of the i-th part; is the i-th enhanced multi-scale feature map.

7. The dual-modal industrial packaging defect detection method according to claim 3, wherein In S34, performing a differentiable binarization operation on the fused feature F to generate a binarization mask of the text area includes: In the formula, k is a hyperparameter that controls the slope of the binarization curve; t is the binarization threshold; B is the binarization mask, representing the text area; e is the natural logarithm base.

8. The dual-modal industrial packaging defect detection method according to claim 1, wherein The specific content of S4 includes: Input the text image corresponding to the text area into the scene text recognition model, convert the text image into character components through hierarchical overlapping patch embedding, and perform three-stage feature extraction on the character components at different scales; Among them, each stage uses a hybrid block for feature extraction: Wherein, is the output feature map of the i-th hybrid block, ; is the input feature map of the i-th hybrid block; is the hybrid block operation; Perform a merging operation on the output feature map as follows: Where C is the merged feature map; is the merging operation; Pass the merged feature map C through a linear prediction layer to recognize each character, compare the recognized characters with the standard text content, and then detect text defects.

9. The dual-modal industrial packaging defect detection method according to claim 1, characterized in that The specific content of S5 includes: Calculate the structural similarity index SSIM between the label area image x1 in the physical label image and the label area image y1 in the label electronic design drawing, and then detect label area image defects: In the formula, and are the average gray values of the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing, respectively; and are the gray variances of the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing, respectively; is the gray covariance of the label area image x1 in the label physical image and the label area image y1 in the label electronic design drawing; and are constants.

Citation Information

Patent Citations

  • Injection product production data management system and method based on block chain

    CN111611607A

  • Filling barrel label character positioning identification method, device and equipment and detection method

    CN116259050A

  • Label text detection method based on multistage separation feature enhancement and spatial attention

    CN116486394A

  • Technology for identifying paper file content and converting paper file content into OFD file with high fidelity

    CN116704540A

  • Model training method and system, seal text recognition method and system, equipment and medium

    CN118196809A

Cited By

  • Display mode self-adaptive switching method based on continuous frame OCR (Optical Character Recognition) analysis

    CN121092259A

  • Corrugated paper box appearance defect detection method based on image recognition

    CN122335829A