Packaging defect detection method and system based on machine vision
The use of machine vision technology to detect packaging defects solves the problem of low efficiency of traditional manual inspection and achieves efficient and accurate packaging defect identification.
Patent Information
- Application Number
- CN202510843273.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional manual visual inspection of packaging defects is inefficient and cannot meet the needs of modern large-scale production. In addition, the consistency and accuracy of the inspection results are difficult to guarantee.
A packaging defect detection method based on machine vision is adopted. The initial image is obtained for preprocessing, the target area image is extracted, and multi-scale feature extraction and feature enhancement fusion are performed. The target recognition model and attention mechanism are used for feature optimization, and finally the defect fault is determined in the database.
It improves the efficiency and accuracy of packaging defect detection, can more accurately identify packaging defects, reduce missed detections and false detections, and improve the consistency and efficiency of detection.
Smart Images

Figure CN120707534A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a packaging defect detection method and system based on machine vision. Background Art
[0002] Packaging, as a crucial component of commodity production and distribution, not only protects goods and facilitates transportation and storage, but also, to a certain extent, influences their appearance and market competitiveness. However, during the packaging process, various defects may develop, such as scratches, stains, and damage to the packaging surface. These defects not only reduce the quality and aesthetics of the packaging but can also cause damage to the goods during transportation and storage, impacting their normal sale and use, and even causing financial losses and reputational damage to the manufacturer.
[0003] Traditional packaging defect detection methods mainly rely on manual visual inspection, which has its shortcomings. On the one hand, manual inspection is inefficient and cannot meet the needs of modern large-scale production. On high-speed production lines, manual inspectors need to observe and judge a large number of packages in a short period of time, which can easily lead to problems such as fatigue and lack of concentration, resulting in missed inspections and false detections. On the other hand, manual inspection is highly subjective, and different inspectors may have different standards for judging defects, which makes it difficult to ensure the consistency and accuracy of the inspection results, resulting in low efficiency of packaging defect detection. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of low efficiency in packaging defect detection and to propose a packaging defect detection method and system based on machine vision.
[0005] In a first aspect of the present invention, a packaging defect detection method based on machine vision is first proposed, the method comprising: Acquiring an initial image, and preprocessing the initial image to obtain an original image; Substituting the original image into the target recognition model to extract the target foreground to obtain a target area map; Continuously downsampling the target area image to obtain a first downsampling image, a second downsampling image, and a third downsampling image; Substituting the third down-sampled image into an enhancement module to obtain a first enhanced image; The first enhanced image is upsampled and the second downsampled Figure 1 Substitute it into the feature synthesis module to obtain the second enhanced image; The second enhanced image is upsampled and then combined with the first downsampled image Figure 1 Substitute it into the feature synthesis module to obtain the third enhanced image; The target area map and the third enhanced image are substituted into an element fusion module to obtain target features, and a database is searched according to the target features to determine the defect fault.
[0006] Optionally, preprocessing the initial image to obtain the original image includes: Performing average pooling and maximum pooling on the initial image to obtain a first average vector and a first maximum vector; Substituting the first average vector and the first maximum vector into a multilayer perceptron to obtain a channel attention map; Multiplying the initial image and the channel attention map to obtain a refined feature map; Performing average pooling and maximum pooling on the refined feature map to obtain a second average vector and a second maximum vector; Concatenate the second average vector and the second maximum vector to obtain a vector concatenation feature, perform a 3×3 convolution operation on the vector concatenation feature, and then perform sigmoid activation on the concatenated feature to obtain a spatial attention map; The spatial attention map and the refined feature map are multiplied to obtain the original image.
[0007] Optionally, substituting the third down-sampled image into an enhancement module to obtain a first enhanced image includes: Substituting the third downsampled graph into the attention mechanism to obtain a query vector, a key vector, and a value vector, performing a dot product operation on the query vector and the key vector to obtain an attention weight matrix, and performing a dot product operation on the attention weight matrix and the value vector to obtain a first feature; performing element-wise addition on the first feature and the third down-sampled image to obtain an intermediate down-sampled image; Substituting the third down-sampled image into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature and the average pooling feature; Performing element-wise multiplication on the intermediate sampling graph according to the maximum pooling feature and the average pooling feature to obtain a maximum intermediate sampling graph and an average intermediate sampling graph; The maximum intermediate sampling image and the average intermediate sampling image are spliced to obtain a first enhanced image.
[0008] Optionally, the workflow of the feature synthesis module includes: Acquire a first input feature and a second input feature, and concatenate the first input feature and the second input feature to obtain a first input concatenated feature; Substituting the first input splicing features into the global feature extraction module and the local feature extraction module respectively to obtain global features and local features; fusing the first input splicing feature, the global feature, and the local feature to obtain a target fused feature; Performing a 1×1 convolution operation on the first input feature and the second input feature respectively to obtain a first input convolution feature and a second input convolution feature; Performing element-wise multiplication on the first input convolution feature and the target fusion feature to obtain a first fusion feature, and performing element-wise multiplication on the second input convolution feature and the target fusion feature to obtain a second fusion feature; The first input splicing feature, the first fusion feature, and the second fusion feature are channel-connected and then a 1×1 convolution operation is performed to obtain the output feature of the feature synthesis module.
[0009] Optionally, substituting the target area map and the third enhanced image into an element fusion module to obtain target features includes: Performing convolution operations on the target area map and the third enhanced image respectively to obtain a first convolution feature and a second convolution feature; Splicing the target area map and the third enhanced image to obtain a first splicing feature; Performing a convolution operation on the first concatenated features and then performing average pooling to obtain a third convolution feature; Performing element-wise multiplication on the first convolution feature and the third convolution feature to obtain a fourth convolution feature, and performing element-wise multiplication on the second convolution feature and the third convolution feature to obtain a fifth convolution feature; The third convolution feature, the fourth convolution feature, and the fifth convolution feature are concatenated to obtain the target feature.
[0010] In a second aspect of the present invention, a packaging defect detection system based on machine vision is proposed, comprising: A preprocessing module is used to obtain an initial image and preprocess the initial image to obtain an original image; A foreground extraction module is used to substitute the original image into the target recognition model to extract the target foreground and obtain a target area map; A downsampling module, configured to continuously downsample the target area image to obtain a first downsampling image, a second downsampling image, and a third downsampling image; A first enhanced image generation module, configured to substitute the third down-sampled image into an enhancement module to obtain a first enhanced image; The second enhanced image generation module is used to upsample the first enhanced image and then perform the downsampling operation on the first enhanced image. Figure 1 Substitute it into the feature synthesis module to obtain the second enhanced image; The third enhanced image generation module is used to upsample the second enhanced image and then perform the upsampling operation on the first downsampled image. Figure 1 Substitute it into the feature synthesis module to obtain the third enhanced image; The target feature determination module is used to substitute the target area map and the third enhanced image into the element fusion module to obtain the target feature, and determine the defect fault according to the target feature.
[0011] Optionally, the preprocessing module includes: A first pooling module, configured to perform average pooling and maximum pooling on the initial image to obtain a first average vector and a first maximum vector; a channel attention map generating module, configured to substitute the first average vector and the first maximum vector into a multilayer perceptron to obtain a channel attention map; A refined feature map generation module, configured to multiply the initial image and the channel attention map to obtain a refined feature map; A second pooling module is used to perform average pooling and maximum pooling on the refined feature map to obtain a second average vector and a second maximum vector; a spatial attention map determining module, configured to concatenate the second average vector and the second maximum vector to obtain a vector concatenation feature, perform a 3×3 convolution operation on the vector concatenation feature, and then perform sigmoid activation on the concatenated feature to obtain a spatial attention map; The original image determination module is used to multiply the spatial attention map and the refined feature map to obtain the original image.
[0012] Optionally, the first enhanced image generation module includes: A first feature generation module is configured to substitute the third downsampled graph into an attention mechanism to obtain a query vector, a key vector, and a value vector, perform a dot product operation on the query vector and the key vector to obtain an attention weight matrix, and perform a dot product operation on the attention weight matrix and the value vector to obtain a first feature; an intermediate sampling map generating module, configured to perform element-wise addition of the first feature and the third down-sampling map to obtain an intermediate sampling map; A third pooling module is used to substitute the third down-sampled image into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature and the average pooling feature; A first element multiplication module, configured to perform element multiplication on the intermediate sampling map according to the maximum pooling feature and the average pooling feature to obtain a maximum intermediate sampling map and an average intermediate sampling map; The sampling image splicing module is used to splice the maximum intermediate sampling image and the average intermediate sampling image to obtain a first enhanced image.
[0013] Optionally, the feature synthesis module includes: A first input splicing feature determination module is configured to obtain a first input feature and a second input feature, and to splice the first input feature and the second input feature to obtain a first input splicing feature; A feature extraction module, configured to substitute the first input splicing features into a global feature extraction module and a local feature extraction module respectively to obtain global features and local features; The target fusion feature determines a first splicing feature module, which is used to fuse the first input splicing feature, the global feature and the local feature to obtain a target fusion feature; An input convolution feature determination module, configured to perform a 1×1 convolution operation on the first input feature and the second input feature to obtain a first input convolution feature and a second input convolution feature; a second fusion feature determination module, configured to perform element-wise multiplication of the first input convolution feature and the target fusion feature to obtain a first fusion feature, and perform element-wise multiplication of the second input convolution feature and the target fusion feature to obtain a second fusion feature; An output feature determination module is used to perform a 1×1 convolution operation on the first input splicing feature, the first fusion feature, and the second fusion feature after channel connection to obtain the output feature of the feature synthesis module.
[0014] Optionally, the target feature determination module includes: A convolution feature extraction module, configured to perform convolution operations on the target area map and the third enhanced image to obtain a first convolution feature and a second convolution feature; a first splicing feature determination module, configured to splice the target area map and the third enhanced image to obtain a first splicing feature; a third convolution feature determination module, configured to perform a convolution operation on the first concatenated features and then perform average pooling to obtain a third convolution feature; a fifth convolution feature determination module, configured to perform element-wise multiplication on the first convolution feature and the third convolution feature to obtain a fourth convolution feature, and perform element-wise multiplication on the second convolution feature and the third convolution feature to obtain a fifth convolution feature; A target feature generation module is used to splice the third convolution feature, the fourth convolution feature and the fifth convolution feature to obtain the target feature.
[0015] Beneficial effects of the present invention: The present invention proposes a packaging defect detection method based on machine vision, which includes obtaining an initial image, preprocessing the initial image to obtain an original image, substituting the original image into a target recognition model to extract the target foreground to obtain a target area image, continuously downsampling the target area image to obtain a first downsampling image, a second downsampling image and a third downsampling image, substituting the third downsampling image into an enhancement module to obtain a first enhanced image, and upsampling the first enhanced image and then combining it with the second downsampling image. Figure 1Substitute the feature synthesis module to obtain the second enhanced image; after upsampling the second enhanced image, combine it with the first downsampling Figure 1 The target region map is then substituted into the feature synthesis module to generate a third enhanced image. The target region map and the third enhanced image are then substituted into the element fusion module to generate target features. Defects are then determined based on the target features. By extracting the foreground from the preprocessed image to generate the target region map, multi-scale feature extraction and feature enhancement fusion are performed on the target region map. The resulting target features can more accurately identify defects, improving the efficiency of packaging defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will be further described below with reference to the accompanying drawings.
[0017] Figure 1 A flowchart of a packaging defect detection method based on machine vision provided by an embodiment of the present invention; Figure 2 A flowchart of another packaging defect detection method based on machine vision provided by an embodiment of the present invention; Figure 3 A framework diagram of a packaging defect detection system based on machine vision provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0019] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0020] The embodiment of the present invention provides a packaging defect detection method based on machine vision. Figure 1 , Figure 1 A flowchart of a packaging defect detection method based on machine vision provided by an embodiment of the present invention. The method includes the following steps: S101, obtaining an initial image, and preprocessing the initial image to obtain an original image; S102, substituting the original image into the target recognition model to extract the target foreground to obtain a target area map; S103, continuously downsampling the target area image to obtain a first downsampling image, a second downsampling image, and a third downsampling image; S104, substituting the third down-sampled image into an enhancement module to obtain a first enhanced image; S105, upsampling the first enhanced image and downsampling the second enhanced image Figure 1Substitute it into the feature synthesis module to obtain the second enhanced image; S106, upsampling the second enhanced image and performing the downsampling operation on the first enhanced image. Figure 1 Substitute it into the feature synthesis module to obtain the third enhanced image; S107 , substituting the target area map and the third enhanced image into an element fusion module to obtain target features, and searching a database to determine the defect fault according to the target features.
[0021] A packaging defect detection method based on machine vision provided by an embodiment of the present invention obtains a target area map by extracting the foreground in the preprocessed image, and then performs multi-scale feature extraction and feature enhancement fusion on the target area map. The obtained target features can more accurately determine defect faults, thereby improving the efficiency of packaging defect detection.
[0022] In one implementation, after obtaining the initial image, preprocessing is performed to obtain the original image, which can effectively remove noise in the image and adjust the brightness of the image. In actual packaging appearance inspection, it is easily affected by external environmental factors, resulting in low image quality. Preprocessing can significantly improve the clarity and recognition of the image, providing high-quality input for subsequent target foreground extraction.
[0023] In one implementation, the original image is substituted into a target recognition model to extract the target foreground and obtain a target area map, which can accurately separate the target from the complex background. The target recognition model used here is a commonly used object recognition model on the market, such as Faster R-CNN, YOLO and other target detection models. The target foreground is the area where the packaging is to be identified, which is the foreground image including the packaging.
[0024] In one implementation, the target area image is continuously downsampled to obtain a first downsampled image, a second downsampled image, and a third downsampled image. Downsampled images of different scales can capture different characteristic information of the target. The small-scale downsampled image (the third downsampled image) can obtain the overall structure and macroscopic features of the target, which helps to identify the overall shape and approximate position of the target; the large-scale downsampled image (the first downsampled image) can retain more detailed information, which is very important for identifying the local features and subtle defects of the target.
[0025] In one implementation, the third downsampled image is substituted into the enhancement module to obtain the first enhanced image. The enhancement module can enhance the third downsampled image to highlight the key information of the target and suppress irrelevant information. The first enhanced image is upsampled and then combined with the second downsampled image. Figure 1 Substitute it into the feature synthesis module to obtain the second enhanced image, and then upsample the second enhanced image and combine it with the first downsampled image. Figure 1The third enhanced image is obtained by substituting it into the feature synthesis module, which realizes the fusion of features of different scales, enabling the model to simultaneously utilize the macro structure and micro detail information of the target, effectively combining the two types of information to improve the accuracy of target recognition.
[0026] In one implementation, the target region map and the third enhanced image are fed into an element fusion module to obtain target features. The element fusion module then comprehensively processes the target region map and the third enhanced image (which has undergone multi-level enhancement and feature synthesis) to extract more representative and discriminative target features. These features can better reflect the essential properties of the target and provide strong support for subsequent defect and fault detection.
[0027] In one implementation, the defect fault is determined by searching the database based on the target features. Since the image is effectively processed and feature extracted in the previous stages, the target features obtained are more accurate and comprehensive, thereby improving the accuracy of defect fault detection; the target features corresponding to different fault types are stored in the database.
[0028] In one embodiment, preprocessing the initial image to obtain the original image includes: S1011, performing average pooling and maximum pooling on the initial image to obtain a first average vector and a first maximum vector; S1012, substituting the first average vector and the first maximum vector into a multilayer perceptron to obtain a channel attention map; S1013, multiplying the initial image and the channel attention map to obtain a refined feature map; S1014, performing average pooling and maximum pooling on the refined feature map to obtain a second average vector and a second maximum vector; S1015, concatenating the second average vector and the second maximum vector to obtain a vector concatenation feature, performing a 3×3 convolution operation on the vector concatenation feature, and then performing sigmoid activation to obtain a spatial attention map; S1016, multiply the spatial attention map and the refined feature map to obtain the original image.
[0029] In one implementation, average pooling and max pooling are performed on the initial image to obtain the first mean vector and the first maximum vector. Average pooling captures global information about a channel, while max pooling highlights the most significant features within a channel. These two vectors are then fed into a multi-layer perceptron to generate a channel attention map. The channel attention map adaptively assigns weights to each channel, highlighting channel features that are more important to the target task and suppressing unimportant ones.
[0030] In one implementation, by introducing a channel attention mechanism, the model can better understand the relationship and importance of different channel features, thereby paying more attention to task-related channels when processing images, improving the model's ability to perceive and utilize channel features. This helps the model to more accurately extract and identify target features in complex scenes, especially in poor lighting conditions.
[0031] In one implementation, a refined feature map is generated by multiplying the initial image and the channel attention map. This operation integrates the channel attention mechanism into the feature map, allowing the features in the refined feature map to focus more on important channel information and remove interference from irrelevant channels, thereby optimizing the feature representation. The refined feature map can more accurately reflect the key information in the image relevant to the target task, providing a better foundation for subsequent spatial attention processing.
[0032] In one implementation, the spatial attention map can adaptively assign weights to different spatial regions in the image, highlighting those regions containing important target information and suppressing interference from background or other irrelevant regions. The spatial attention mechanism enables the model to better understand the spatial structure and layout of the image, focusing on the specific position and shape information of the target in the image, which can improve the model's ability to utilize spatial information and its depth of understanding of image content.
[0033] In one implementation, the spatial attention mechanism is further integrated into the feature map, resulting in a feature map that not only includes channel features optimized by channel attention but also highlights important spatial region features. This dual attention mechanism can further improve feature quality, making the features more representative and discriminative, and supporting subsequent image processing tasks.
[0034] In one implementation, the entire process can adaptively extract and optimize image features through the organic combination of channel attention and spatial attention mechanisms, thereby improving the model's ability to understand and process images.
[0035] In one embodiment, substituting the third downsampled image into the enhancement module to obtain the first enhanced image includes: Substitute the third downsampled graph into the attention mechanism to obtain the query vector, key vector and value vector, perform a dot product operation on the query vector and the key vector to obtain the attention weight matrix, and perform a dot product operation on the attention weight matrix and the value vector to obtain the first feature; Performing element-wise addition on the first feature and the third downsampled image to obtain an intermediate downsampled image; Substitute the third downsampled image into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature and the average pooling feature; Multiply the intermediate sampling map elements by the maximum pooling feature and the average pooling feature to obtain the maximum intermediate sampling map and the average intermediate sampling map; The maximum intermediate sampling image and the average intermediate sampling image are spliced to obtain a first enhanced image.
[0036] In one implementation, the attention weight matrix is obtained by performing a dot product operation on the query vector and the key vector, and then the first feature is obtained by performing a dot product operation on the value vector. This can adaptively focus on features at different locations in the third downsampled image, assigning attention weights based on the correlation between features, thereby accurately capturing the key features that are more important to the target task. The attention mechanism dynamically adjusts feature weights, allowing the first feature to better reflect the semantic relevance and importance between different regions in the image, enhancing the expressiveness and discriminability of the features. Compared to traditional convolution operations, the attention mechanism can more flexibly handle features of different scales and locations, improving the model's ability to model complex image features.
[0037] In one implementation, the first feature and the third down-sampled image are element-wise added to obtain an intermediate sampling image. This operation fuses the first feature extracted by the attention mechanism with the original third down-sampled image, retaining the basic feature information in the original image while incorporating key features optimized by the attention mechanism. It fully utilizes the advantages of the original image and attention features, avoids information loss, and gives the intermediate sampling image a richer feature representation.
[0038] In one implementation, the intermediate sampling map is element-wise multiplied according to the maximum pooling feature and the average pooling feature to obtain the maximum intermediate sampling map and the average intermediate sampling map, so that the features in the intermediate sampling map interact with the pooling feature. The features of the intermediate sampling map are further screened and strengthened according to the different characteristics of the pooling feature, thereby enhancing the pertinence and effectiveness of the features.
[0039] In one implementation, multi-level feature extraction and optimization are performed on the third downsampled image through various methods, including attention mechanisms, pooling operations, and feature fusion. The resulting first enhanced image has stronger feature representation capabilities and higher information content. This helps improve the model's performance in various image processing tasks, such as improving the accuracy of image fault detail detection.
[0040] In one embodiment, the workflow of the feature synthesis module includes: Acquire a first input feature and a second input feature, and concatenate the first input feature and the second input feature to obtain a first input concatenated feature; Substituting the first input splicing feature into the global feature extraction module and the local feature extraction module respectively to obtain the global feature and the local feature; Fusing the first input splicing feature, the global feature and the local feature to obtain the target fusion feature; Performing a 1×1 convolution operation on the first input feature and the second input feature respectively to obtain the first input convolution feature and the second input convolution feature; Perform element-wise multiplication on the first input convolution feature and the target fusion feature to obtain a first fusion feature, and perform element-wise multiplication on the second input convolution feature and the target fusion feature to obtain a second fusion feature; The first input splicing feature, the first fusion feature and the second fusion feature are channel-connected and then a 1×1 convolution operation is performed to obtain the output feature of the feature synthesis module.
[0041] In one implementation, the first input feature and the second input feature are spliced to obtain a first input spliced feature. This operation integrates features from different sources so that the model can use multiple types of information at the same time. The spliced features contain richer information, which helps to improve the model's ability to understand the image. When the first input feature is the first enhanced image after upsampling, the second input feature is the second downsampling image. Similarly, other first input features and second input features are as above.
[0042] In one implementation, the first input concatenated features, global features, and local features are fused to create a target fused feature. This process organically combines different types of features, leveraging their respective strengths. The target fused feature incorporates both the original input features and key elements of global and local features, offering greater expressiveness and discrimination, and better reflecting the essential characteristics of the image.
[0043] In one implementation, a 1×1 convolution can linearly transform the channel dimension of a feature map without changing its spatial size, enabling information interaction and fusion between channels. It can also reduce or increase dimensionality, reducing computational effort or increasing feature diversity. The element-wise multiplication operation is equivalent to weighting features, with the target fusion feature acting as the weight to adjust the first and second input convolutional features. This can highlight features more relevant to the target task, suppress irrelevant features, and improve feature effectiveness and relevance.
[0044] In one implementation, the first input concatenated feature, the first fused feature, and the second fused feature are channel-connected to integrate the features processed at different stages to form a more comprehensive feature representation. The channel connection operation retains all the information of each feature, enabling the model to fully utilize the advantages of processing at different stages.
[0045] In one embodiment, substituting the target region map and the third enhanced image into the element fusion module to obtain the target features includes: Performing convolution operations on the target area map and the third enhanced image respectively to obtain a first convolution feature and a second convolution feature; splicing the target area map and the third enhanced image to obtain a first splicing feature; Perform convolution operation on the first concatenated feature and then perform average pooling to obtain the third convolution feature; Performing element-wise multiplication on the first convolution feature and the third convolution feature to obtain a fourth convolution feature, and performing element-wise multiplication on the second convolution feature and the third convolution feature to obtain a fifth convolution feature; The third convolution feature, the fourth convolution feature and the fifth convolution feature are concatenated to obtain the target feature.
[0046] In one implementation, convolution operations are performed on the target area map and the third enhanced image to generate first and second convolution features, respectively. This adjusts the feature dimensions and makes them more suitable for subsequent processing steps. The target area map and the third enhanced image are then concatenated to generate first concatenated features, integrating features from two different sources or processing stages. The target area map contains the original target information, while the third enhanced image has been enhanced to more prominently represent the target features. This concatenation operation enables the model to utilize both types of information simultaneously, enriching the feature representation and helping to improve the model's ability to recognize and understand targets.
[0047] In one implementation, convolution operations can further abstract and extract the concatenated features, capturing higher-level semantic features. Average pooling can reduce the spatial dimensionality of features, reducing computational effort while preserving important feature information. Average pooling averages the feature values within a local region, imbuing the features with a certain degree of translation invariance, improving the model's robustness to changes in target position.
[0048] In one implementation, the first convolution feature and the third convolution feature are element-wise multiplied to obtain the fourth convolution feature, and the second convolution feature and the third convolution feature are element-wise multiplied to obtain the fifth convolution feature. Since the third convolution feature has undergone convolution and pooling operations and contains more advanced semantic information and global features, element-wise multiplication can highlight the key features of the first and second convolution features that are more relevant to the target task, suppress irrelevant features, and improve the effectiveness and pertinence of the features.
[0049] In one implementation method, the third convolution feature, the fourth convolution feature, and the fifth convolution feature are spliced together to obtain the target feature, and the features that have been processed and fused differently are integrated together to form a more comprehensive and richer feature representation. The third convolution feature contains high-level semantic features that have been pooled, and the fourth convolution feature and the fifth convolution feature highlight the key features in the target area map and the third enhanced image, respectively. Through splicing, the target feature can comprehensively utilize this information to better reflect the essential characteristics of the target and provide stronger support for subsequent classification, detection and other tasks.
[0050] In one implementation, multiple operations are performed on the target area map and the third enhanced image to extract features, fuse, and optimize them, so that more representative and discriminative target features can be extracted, thereby improving the accuracy of the model in tasks such as target recognition and defect detection.
[0051] Based on the same inventive concept, the present invention also provides a packaging defect detection system based on machine vision. Figure 2 , Figure 2 A framework diagram of a packaging defect detection system based on machine vision provided by an embodiment of the present invention includes: A preprocessing module is used to obtain an initial image and preprocess the initial image to obtain an original image; The foreground extraction module is used to substitute the original image into the target recognition model to extract the target foreground and obtain the target area map; A downsampling module, configured to continuously downsample the target area image to obtain a first downsampling image, a second downsampling image, and a third downsampling image; A first enhanced image generation module, configured to substitute the third down-sampled image into the enhancement module to obtain a first enhanced image; The second enhanced image generation module is used to upsample the first enhanced image and then downsample it Figure 1 Substitute it into the feature synthesis module to obtain the second enhanced image; The third enhanced image generation module is used to upsample the second enhanced image and then combine it with the first downsampled image. Figure 1 Substitute it into the feature synthesis module to obtain the third enhanced image; The target feature determination module is used to substitute the target area map and the third enhanced image into the element fusion module to obtain the target feature, and determine the defect fault according to the target feature search database.
[0052] A packaging defect detection system based on machine vision provided by an embodiment of the present invention obtains a target area map by extracting the foreground in the preprocessed image, and then performs multi-scale feature extraction and feature enhancement fusion on the target area map. The obtained target features can more accurately determine defect faults, thereby improving the efficiency of packaging defect detection.
[0053] In one embodiment, the pre-processing module includes: A first pooling module, configured to perform average pooling and maximum pooling on the initial image to obtain a first average vector and a first maximum vector; a channel attention map generation module, configured to substitute the first average vector and the first maximum vector into a multilayer perceptron to obtain a channel attention map; The refined feature map generation module is used to multiply the initial image and the channel attention map to obtain the refined feature map; A second pooling module is used to perform average pooling and maximum pooling on the refined feature map to obtain a second average vector and a second maximum vector; A spatial attention map determination module is used to concatenate the second average vector and the second maximum vector to obtain a vector concatenation feature, perform a 3×3 convolution operation on the vector concatenation feature, and then input a sigmoid activation to obtain a spatial attention map; The original image determination module is used to multiply the spatial attention map and the refined feature map to obtain the original image.
[0054] In one embodiment, the first enhanced image generation module includes: A first feature generation module is configured to substitute the third downsampled graph into the attention mechanism to obtain a query vector, a key vector, and a value vector, perform a dot product operation on the query vector and the key vector to obtain an attention weight matrix, and perform a dot product operation on the attention weight matrix and the value vector to obtain a first feature; an intermediate sampling map generation module, configured to perform element-wise addition of the first feature and the third down-sampling map to obtain an intermediate sampling map; The third pooling module is used to substitute the third down-sampled image into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature and the average pooling feature; The first element multiplication module is used to perform element multiplication on the intermediate sampling map according to the maximum pooling feature and the average pooling feature to obtain the maximum intermediate sampling map and the average intermediate sampling map; The sampling image stitching module is used to stitch the maximum intermediate sampling image and the average intermediate sampling image to obtain a first enhanced image.
[0055] In one embodiment, the feature synthesis module includes: A first input splicing feature determination module is used to obtain a first input feature and a second input feature, and splice the first input feature and the second input feature to obtain a first input splicing feature; A feature extraction module, configured to substitute the first input splicing feature into a global feature extraction module and a local feature extraction module respectively to obtain a global feature and a local feature; The target fusion feature determines a first splicing feature module, which is used to fuse the first input splicing feature, the global feature and the local feature to obtain the target fusion feature; An input convolution feature determination module, configured to perform a 1×1 convolution operation on the first input feature and the second input feature respectively to obtain a first input convolution feature and a second input convolution feature; a second fusion feature determination module, configured to perform element-wise multiplication of the first input convolution feature and the target fusion feature to obtain a first fusion feature, and perform element-wise multiplication of the second input convolution feature and the target fusion feature to obtain a second fusion feature; The output feature determination module is used to perform a 1×1 convolution operation on the first input splicing feature, the first fusion feature, and the second fusion feature after channel connection to obtain the output feature of the feature synthesis module.
[0056] In one embodiment, the target feature determination module includes: A convolution feature extraction module, configured to perform convolution operations on the target area map and the third enhanced image to obtain a first convolution feature and a second convolution feature; a first splicing feature determination module, configured to splice the target area image and the third enhanced image to obtain a first splicing feature; A third convolution feature determination module is used to perform a convolution operation on the first concatenated feature and then perform average pooling to obtain a third convolution feature; a fifth convolution feature determination module, configured to perform element-wise multiplication on the first convolution feature and the third convolution feature to obtain a fourth convolution feature, and perform element-wise multiplication on the second convolution feature and the third convolution feature to obtain a fifth convolution feature; The target feature generation module is used to concatenate the third convolution feature, the fourth convolution feature and the fifth convolution feature to obtain the target feature.
[0057] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A packaging defect detection method based on machine vision, characterized in that: The method comprises: Acquiring an initial image, and preprocessing the initial image to obtain an original image; Substituting the original image into the target recognition model to extract the target foreground to obtain a target area map; Continuously downsampling the target area image to obtain a first downsampling image, a second downsampling image, and a third downsampling image; Substituting the third down-sampled image into an enhancement module to obtain a first enhanced image; Upsampling the first enhanced image and substituting it together with the second downsampled image into a feature synthesis module to obtain a second enhanced image; Upsampling the second enhanced image and substituting it together with the first downsampled image into a feature synthesis module to obtain a third enhanced image; The target area map and the third enhanced image are substituted into an element fusion module to obtain target features, and a database is searched according to the target features to determine the defect fault.
2. The packaging defect detection method based on machine vision according to claim 1, characterized in that: Preprocessing the initial image to obtain the original image includes: Performing average pooling and maximum pooling on the initial image to obtain a first average vector and a first maximum vector; Substituting the first average vector and the first maximum vector into a multilayer perceptron to obtain a channel attention map; Multiplying the initial image and the channel attention map to obtain a refined feature map; Performing average pooling and maximum pooling on the refined feature map to obtain a second average vector and a second maximum vector; Concatenate the second average vector and the second maximum vector to obtain a vector concatenation feature, perform a 3×3 convolution operation on the vector concatenation feature, and then perform sigmoid activation on the concatenated feature to obtain a spatial attention map; The spatial attention map and the refined feature map are multiplied to obtain the original image.
3. The packaging defect detection method based on machine vision according to claim 1, characterized in that: Substituting the third down-sampled image into the enhancement module to obtain a first enhanced image comprises: Substituting the third downsampled graph into the attention mechanism to obtain a query vector, a key vector, and a value vector, performing a dot product operation on the query vector and the key vector to obtain an attention weight matrix, and performing a dot product operation on the attention weight matrix and the value vector to obtain a first feature; performing element-wise addition on the first feature and the third down-sampled image to obtain an intermediate down-sampled image; Substituting the third down-sampled image into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature and the average pooling feature; Performing element-wise multiplication on the intermediate sampling graph according to the maximum pooling feature and the average pooling feature to obtain a maximum intermediate sampling graph and an average intermediate sampling graph; The maximum intermediate sampling image and the average intermediate sampling image are spliced to obtain a first enhanced image.
4. The packaging defect detection method based on machine vision according to claim 1, characterized in that: The workflow of the feature synthesis module includes: Acquire a first input feature and a second input feature, and concatenate the first input feature and the second input feature to obtain a first input concatenated feature; Substituting the first input splicing features into the global feature extraction module and the local feature extraction module respectively to obtain global features and local features; fusing the first input splicing feature, the global feature, and the local feature to obtain a target fused feature; Performing a 1×1 convolution operation on the first input feature and the second input feature respectively to obtain a first input convolution feature and a second input convolution feature; Performing element-wise multiplication on the first input convolution feature and the target fusion feature to obtain a first fusion feature, and performing element-wise multiplication on the second input convolution feature and the target fusion feature to obtain a second fusion feature; The first input splicing feature, the first fusion feature, and the second fusion feature are channel-connected and then a 1×1 convolution operation is performed to obtain the output feature of the feature synthesis module.
5. The packaging defect detection method based on machine vision according to claim 1, characterized in that: Substituting the target area map and the third enhanced image into an element fusion module to obtain target features includes: Performing convolution operations on the target area map and the third enhanced image respectively to obtain a first convolution feature and a second convolution feature; Splicing the target area map and the third enhanced image to obtain a first splicing feature; Performing a convolution operation on the first concatenated features and then performing average pooling to obtain a third convolution feature; Performing element-wise multiplication on the first convolution feature and the third convolution feature to obtain a fourth convolution feature, and performing element-wise multiplication on the second convolution feature and the third convolution feature to obtain a fifth convolution feature; The third convolution feature, the fourth convolution feature, and the fifth convolution feature are concatenated to obtain the target feature.
6. A packaging defect detection system based on machine vision, characterized in that: The system comprises: A preprocessing module is used to obtain an initial image and preprocess the initial image to obtain an original image; A foreground extraction module is used to substitute the original image into the target recognition model to extract the target foreground and obtain a target area map; A downsampling module, configured to continuously downsample the target area image to obtain a first downsampling image, a second downsampling image, and a third downsampling image; A first enhanced image generation module, configured to substitute the third down-sampled image into an enhancement module to obtain a first enhanced image; A second enhanced image generation module is configured to upsample the first enhanced image and substitute the upsampled image together with the second downsampled image into a feature synthesis module to obtain a second enhanced image; a third enhanced image generation module, configured to upsample the second enhanced image and substitute the upsampled image together with the first downsampled image into a feature synthesis module to obtain a third enhanced image; The target feature determination module is used to substitute the target area map and the third enhanced image into the element fusion module to obtain the target feature, and determine the defect fault according to the target feature.
7. The packaging defect detection system based on machine vision according to claim 6, characterized in that: The pre-processing module comprises: A first pooling module, configured to perform average pooling and maximum pooling on the initial image to obtain a first average vector and a first maximum vector; a channel attention map generating module, configured to substitute the first average vector and the first maximum vector into a multilayer perceptron to obtain a channel attention map; A refined feature map generation module, configured to multiply the initial image and the channel attention map to obtain a refined feature map; A second pooling module is used to perform average pooling and maximum pooling on the refined feature map to obtain a second average vector and a second maximum vector; a spatial attention map determining module, configured to concatenate the second average vector and the second maximum vector to obtain a vector concatenation feature, perform a 3×3 convolution operation on the vector concatenation feature, and then perform sigmoid activation on the concatenated feature to obtain a spatial attention map; The original image determination module is used to multiply the spatial attention map and the refined feature map to obtain the original image.
8. The packaging defect detection system based on machine vision according to claim 6, characterized in that: The first enhanced image generation module includes: A first feature generation module is configured to substitute the third downsampled graph into an attention mechanism to obtain a query vector, a key vector, and a value vector, perform a dot product operation on the query vector and the key vector to obtain an attention weight matrix, and perform a dot product operation on the attention weight matrix and the value vector to obtain a first feature; an intermediate sampling map generating module, configured to perform element-wise addition of the first feature and the third down-sampling map to obtain an intermediate sampling map; A third pooling module is used to substitute the third down-sampled image into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling feature and the average pooling feature; A first element multiplication module, configured to perform element multiplication on the intermediate sampling map according to the maximum pooling feature and the average pooling feature to obtain a maximum intermediate sampling map and an average intermediate sampling map; The sampling image splicing module is used to splice the maximum intermediate sampling image and the average intermediate sampling image to obtain a first enhanced image.
9. The packaging defect detection system based on machine vision according to claim 6, characterized in that: The feature synthesis module includes: A first input splicing feature determination module is configured to obtain a first input feature and a second input feature, and to splice the first input feature and the second input feature to obtain a first input splicing feature; A feature extraction module, configured to substitute the first input splicing features into a global feature extraction module and a local feature extraction module respectively to obtain global features and local features; The target fusion feature determines a first splicing feature module, which is used to fuse the first input splicing feature, the global feature and the local feature to obtain a target fusion feature; An input convolution feature determination module, configured to perform a 1×1 convolution operation on the first input feature and the second input feature to obtain a first input convolution feature and a second input convolution feature; a second fusion feature determination module, configured to perform element-wise multiplication of the first input convolution feature and the target fusion feature to obtain a first fusion feature, and perform element-wise multiplication of the second input convolution feature and the target fusion feature to obtain a second fusion feature; An output feature determination module is used to perform a 1×1 convolution operation on the first input splicing feature, the first fusion feature, and the second fusion feature after channel connection to obtain the output feature of the feature synthesis module.
10. The packaging defect detection system based on machine vision according to claim 6, characterized in that: The target feature determination module includes: A convolution feature extraction module, configured to perform convolution operations on the target area map and the third enhanced image to obtain a first convolution feature and a second convolution feature; a first splicing feature determination module, configured to splice the target area map and the third enhanced image to obtain a first splicing feature; a third convolution feature determination module, configured to perform a convolution operation on the first concatenated features and then perform average pooling to obtain a third convolution feature; a fifth convolution feature determination module, configured to perform element-wise multiplication on the first convolution feature and the third convolution feature to obtain a fourth convolution feature, and perform element-wise multiplication on the second convolution feature and the third convolution feature to obtain a fifth convolution feature; A target feature generation module is used to splice the third convolution feature, the fourth convolution feature and the fifth convolution feature to obtain the target feature.