Cigarette packaging paper surface defect detection method based on double-domain attention

By enhancing and fusing multi-scale feature maps through a dual-domain attention mechanism, the problems of low efficiency and insufficient information fusion in cigarette packaging paper inspection are solved, and high-precision defect detection and positioning are achieved.

CN120707559APending Publication Date: 2025-09-26CHINA TOBACCO HENAN IND CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511133200.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing cigarette packaging paper detection methods are inefficient, costly, easily affected by subjective factors, and difficult to achieve full coverage detection. In addition, existing image detection methods fail to effectively integrate multi-scale information, resulting in missed detections or false detections.

Method used

A cigarette packaging paper surface defect detection method based on dual-domain attention is adopted. The multi-scale feature maps are enhanced and fused through the dual-domain attention mechanism, the frequency domain and spatial domain information are integrated, cross-scale refinement processing is performed, and the image is reconstructed to determine defects.

Benefits of technology

It significantly improves the accuracy of cigarette packaging paper defect detection and the ability to locate printing defect edges, solves the problems of missing information and blurred details, and meets the needs of high-precision automated quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707559A_ABST
    Figure CN120707559A_ABST
Patent Text Reader

Abstract

The invention discloses a cigarette packaging paper surface defect detection method based on double-domain attention. The method mainly comprises the following steps: extracting a multi-scale feature map of a to-be-detected cigarette packaging paper image; enhancing the area where the defect is located by using double-domain attention; performing multi-scale feature fusion on the enhanced deeper feature map, and integrating context information to obtain an enhanced feature map; performing cross-scale refining processing on the enhanced feature map to obtain a refined feature map; splicing the enhanced first-layer feature map and the refined feature map to obtain a reconstructed image; and judging whether surface defects and defect positions exist in the to-be-detected cigarette packaging paper image or not by utilizing the reconstructed image. According to the method, frequency domain and space domain information is integrated, multi-scale information interaction is realized, the defect detection accuracy of the cigarette packaging paper is remarkably improved, particularly, the detection precision of the printing defects on the surface of the cigarette packaging paper and the edge positioning capability of the printing defects are improved, and a series of problems in the conventional detection scheme of the surface defects of the cigarette packaging paper are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cigarette manufacturing, and in particular to a method for detecting surface defects of cigarette packaging paper based on dual-domain attention. Background Art

[0002] In the tobacco industry, the printing quality of cigarette wrappers (specifically, small packs, carton packs, etc.) is a core guarantee of brand value and consumer trust. Traditional inspection relies on manual visual inspection, where quality inspectors visually check each cigarette wrapper for pattern integrity, color consistency, and surface defects (such as scratches, missing prints, and misalignment). However, manual inspection is inefficient, costly, and susceptible to subjective judgment. Full coverage is particularly difficult to achieve on high-speed production lines, and the risk of missed and false detections is significant. Furthermore, manual standards are difficult to standardize, and differences in operator experience can easily lead to volatile inspection results, making it difficult to meet the urgent demand for high-precision, automated quality inspection in modern industry.

[0003] With the iteration of technology, the industry has gradually tried to adopt image detection methods such as computer vision to replace traditional inspection methods. However, how to make full use of image information and guide the model to focus on key feature areas is a major technical challenge.

[0004] Existing methods often focus on processing spatial information. While they can capture pixel-level details, they overlook the importance of frequency domain information. Another core challenge lies in the upsampling operation during feature reconstruction. Traditional upsampling methods are prone to losing details or introducing artifacts when recovering high-resolution features due to the smoothness of the interpolation algorithm itself, severely impairing the clarity and localization accuracy of defect edges. Furthermore, existing methods often employ a single-scale upsampling strategy and lack a collaborative optimization mechanism for cross-level features. Deep features contain rich semantic information but have low spatial resolution, while shallow features retain detail but are subject to high noise interference. Without effective fusion of multi-scale information, the reconstructed feature map may not be able to simultaneously account for both the global semantics and local details of cigarette packaging defects. For example, in cigarette packaging defect detection scenarios, large-area underprinting defects require deep semantic analysis. Existing upsampling techniques struggle to balance these two requirements, resulting in detection results that are prone to missed or false detections. Summary of the Invention

[0005] In view of the above, the present invention aims to provide a cigarette packaging paper surface defect detection method based on dual-domain attention to solve the above-mentioned technical problems.

[0006] The technical solution adopted in the present invention is as follows:

[0007] The present invention provides a method for detecting surface defects of cigarette packaging paper based on dual-domain attention, which includes:

[0008] Perform feature extraction on the original image of the cigarette packaging paper to be tested to obtain a multi-scale feature map;

[0009] Using dual-domain attention to enhance the region of interest of the multi-scale feature map, the region of interest is the area where the surface defects of the cigarette packaging paper are located;

[0010] The enhanced feature maps of other layers except the first layer are subjected to multi-scale feature fusion processing, and context information is integrated to obtain multi-scale context interaction enhanced feature maps;

[0011] Perform cross-scale refinement processing on the multi-scale context interaction enhanced feature map to obtain a refined feature map;

[0012] The first-layer feature map enhanced by dual-domain attention is concatenated with the refined feature map to obtain a reconstructed image.

[0013] The reconstructed image is used to determine whether there are surface defects and the location of the defects in the original image of the cigarette packaging paper to be tested.

[0014] In at least one possible implementation, the dual-domain attention enhancement method for the multi-scale feature map includes:

[0015] The multi-scale feature map is input into the wavelet attention enhancement network, and the feature map is decoupled by discrete wavelet transform to generate the frequency domain enhanced feature map;

[0016] The multi-scale feature map is input into the spatial attention enhancement network, and the spatial saliency weight is extracted through the global pooling operation to generate the spatially enhanced feature map;

[0017] The frequency domain enhanced feature map and the spatial domain enhanced feature map are added pixel by pixel to obtain the dual-domain attention enhanced feature map.

[0018] In at least one possible implementation, generating the frequency-domain enhanced feature map includes:

[0019] The input multi-scale feature map is decoupled into four frequency sub-band feature maps through discrete wavelet transform, including low-frequency sub-band, horizontal high-frequency sub-band, vertical high-frequency sub-band and diagonal high-frequency sub-band;

[0020] The horizontal, vertical and diagonal high-frequency sub-bands are added element by element, and the frequency domain attention map is generated through the Sigmoid activation function;

[0021] Multiplying the frequency domain attention map by the low-frequency subband feature map pixel by pixel to obtain a wavelet enhanced feature map;

[0022] The wavelet enhanced feature map is concatenated with the original input multi-scale feature map along the channel dimension to obtain a frequency domain enhanced feature map.

[0023] In at least one possible implementation, generating a spatially enhanced feature map includes:

[0024] Perform global average pooling and global maximum pooling operations on the input multi-scale feature map to obtain two pooled feature maps;

[0025] The two pooled feature maps are concatenated along the channel dimension, and a spatial attention weight map is generated through a convolutional layer and a Sigmoid activation function;

[0026] The spatial attention weight map is multiplied pixel by pixel with the multi-scale feature map of the original input and the feature response strength is adjusted to obtain a spatially enhanced feature map.

[0027] In at least one possible implementation manner, obtaining a multi-scale context interaction enhanced feature map includes:

[0028] The deep feature map enhanced by dual-domain attention is used to interact with contextual information through a progressive cross-scale feature aggregation mechanism;

[0029] Perform single-branch convolution on shallow features to generate preliminary fusion features;

[0030] After resizing the initial fusion features, they are fused with the next layer features through a two-branch convolution;

[0031] After the previous fusion features are scaled and matched with the current layer features, they are subjected to multi-branch convolution processing with the current layer features and aggregated layer by layer;

[0032] Generate multi-scale contextual interaction enhanced feature maps for adaptive fusion of shallow details and deep semantics.

[0033] In at least one possible implementation manner, obtaining the refined feature map includes:

[0034] Divide the multi-scale context interaction enhancement feature map into shallow group feature maps and deep group feature maps;

[0035] Each set of feature maps is fed into a cross-scale wavelet-refined cross-attention network;

[0036] The discrete wavelet transform is used to perform feature decoupling and lossless downsampling on each group of feature maps to obtain multiple decomposed wavelet sub-bands. The multi-head cross-attention mechanism is combined to interact information between feature maps of different scales and output the refined feature maps.

[0037] In at least one possible implementation manner, obtaining the reconstructed image includes:

[0038] The two sets of feature maps after cross-scale wavelet refinement are upsampled separately to match the size of the first layer feature map after dual-domain attention enhancement;

[0039] The upsampled refined feature map is concatenated with the first-layer attention-enhanced feature map in the channel dimension to form an integrated feature map.

[0040] The integrated feature map is upsampled again to restore it to its original size, and finally the reconstructed feature map is obtained.

[0041] In at least one possible implementation, the feature extraction includes: using the backbone network of a pre-built cigarette packaging paper detection model to extract features from the original image, and obtaining a multi-scale feature map whose spatial size decreases successively from the first layer to the last layer and whose number of channels increases successively.

[0042] Compared with the existing technology, the main design concept of the present invention is to extract a multi-scale feature map of the cigarette packaging paper image to be tested; use dual-domain attention to enhance the area where the surface defects are located in the multi-scale feature map; perform multi-scale feature fusion processing on the enhanced deeper feature map and integrate context information to obtain a multi-scale context interaction enhanced feature map; perform cross-scale refinement processing on the multi-scale context interaction enhanced feature map to obtain a refined feature map; splice the enhanced first-layer feature map with the refined feature map to obtain a reconstructed image; and use the reconstructed image to determine whether there are surface defects and the defect location in the cigarette packaging paper image to be tested. The present invention integrates frequency domain and spatial domain information to achieve multi-scale information interaction, significantly improving the accuracy of cigarette packaging paper defect detection, especially improving the detection accuracy of cigarette packaging paper surface printing defects and the ability to locate printing defect edges, solving the problems of information loss, blurred details and semantic gaps in previous cigarette packaging paper surface defect detection solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:

[0044] Figure 1 A schematic flow chart of a method for detecting surface defects of cigarette packaging paper based on dual-domain attention provided by an embodiment of the present invention;

[0045] Figure 2 This is an architecture diagram of a cigarette packaging paper surface defect detection model based on dual-domain attention provided by an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of dual-domain attention enhancement provided by an embodiment of the present invention;

[0047] Figure 4 A schematic diagram of multi-scale feature fusion provided by an embodiment of the present invention;

[0048] Figure 5 A schematic diagram of cross-scale wavelet refinement provided by an embodiment of the present invention;

[0049] Figure 6 A schematic diagram of discrete wavelet transform provided by an embodiment of the present invention;

[0050] Figure 7 This is a schematic diagram of an example of detecting printing defects on the surface of cigarette packaging paper provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0052] The present invention proposes an embodiment of a method for detecting surface defects of cigarette packaging paper based on dual-domain attention. Specifically, Figure 1 and Figure 2 shown, including:

[0053] Step S1, extracting features from the original image of the cigarette packaging paper to be tested to obtain a multi-scale feature map;

[0054] This step can be expanded upon as follows: the original image data of the cigarette packaging paper to be tested is subjected to feature extraction through the backbone network to obtain the first to fifth layers of multi-scale feature maps. It can be understood that the spatial size of the first to fifth layers of multi-scale feature maps decreases layer by layer, and the number of channels increases layer by layer. Taking the above five-layer feature map as an example, the spatial scale of the first to fifth layers is halved layer by layer, and the number of channels increases layer by layer. The larger the spatial scale, the more detailed information it contains, which is crucial for accurately locating the outline of printing defects; the deeper feature maps with smaller spatial scales mainly contain semantic information, which helps the feature extraction network learn higher-level features.

[0055] In a specific embodiment, EfficientNet is used to extract features from the original image of the cigarette wrapper to be tested, obtaining multi-scale feature maps of different sizes and numbers of channels from the first to fifth layers. EfficientNet's multi-scale feature extraction capability enables it to capture details of the printed pattern on the cigarette wrapper at different scales. During the feature extraction process, multi-scale feature maps of different sizes and numbers of channels are extracted. These feature maps cover information at all levels, from local subtle features such as ink particle distribution to the overall pattern outline. This facilitates a comprehensive analysis of various defects that may occur in the printing of cigarette wrappers, including but not limited to ink detachment, color deviation, blurred patterns, missing characters, etc., providing a solid foundation for subsequent defect detection and location.

[0056] This step also corresponds to the training process of the defect detection model. Specifically, the original images of the cigarette wrappers to be tested are pre-collected surface images of cigarette wrappers without and with printing defects. Before inputting them into the defect detection model for training, the surface images of cigarette wrappers with printing defects are annotated with the printing defects to obtain a binary labeled image corresponding to the original image. White pixels in the binary image represent areas with printing defects, while black pixels represent areas without printing defects.

[0057] Step S2: using dual-domain attention to enhance the region of interest of the multi-scale feature map, where the region of interest is the area where the surface defects of the cigarette packaging paper are located;

[0058] Combined with the previous example, the dual-domain attention mechanism is used to enhance the regions of interest of the first to fifth layers of multi-scale feature maps. The regions of interest are the areas where defects on the surface of cigarette packaging paper are located. That is, the dual-domain attention mechanism is used to enhance the feature maps, strengthening the high-frequency details and spatial key areas related to printing defects on the surface of cigarette packaging paper from the frequency domain and spatial domain respectively.

[0059] To elaborate, in a specific embodiment, as shown in formula (1), the feature map is first processed as a whole using discrete wavelet transform, decomposing it into feature maps of four different frequency subbands, namely low-frequency subband, horizontal high-frequency subband, vertical high-frequency subband, and diagonal high-frequency subband. Among them, the low-frequency subband contains most of the low-frequency information of the original feature map and reflects the overall structure of the image. The horizontal high-frequency subband, vertical high-frequency subband, and diagonal high-frequency subband capture the high-frequency details of the image in the horizontal, vertical, and diagonal directions, respectively. These high-frequency details are often closely related to key features such as the edges and textures of printing defects. In order to give full play to the sensitivity of the high-frequency subband to defect features, the horizontal, vertical, and diagonal high-frequency subbands are added pixel by pixel to integrate high-frequency feature information in multiple directions. Then, the obtained feature map is processed by the Sigmoid activation function to generate an attention map. The attention map is multiplied pixel by pixel with the original low-frequency subband feature map to obtain a wavelet attention enhancement map. This operation enables the high-frequency details related to potential defects to be enhanced in a targeted manner based on the low-frequency structure, thereby obtaining a richer and more prominent feature representation.

[0060]

[0061] in, Represents the multi-scale feature map of the input, represents the frequency domain attention map, represents the frequency domain enhanced feature map, represents upsampling, represents element-by-element addition, Represents element-wise multiplication.

[0062] Furthermore, as shown in formula (2), spatial attention includes: The spatial attention mechanism enhances the feature map from another perspective. The input feature map is processed using average pooling and maximum pooling respectively. Average pooling can capture the overall average response in the feature map, while maximum pooling can highlight local significant features. By splicing the two pooling results, a feature representation that integrates global and local information is obtained. The spliced ​​feature map is processed by the Sigmoid activation function to obtain a spatial attention map, which emphasizes the more important areas in the spatial position of the feature map. Finally, the spatial attention map is multiplied pixel by pixel with the original feature map to obtain a spatial attention enhancement map, which further strengthens the discriminative spatial information in the feature map.

[0063]

[0064] in Represents the position in the input feature map The value of is the size of the pooling window, Position in the pooled output feature map The value of represents the spatial attention map, represents the spatial domain enhanced feature map, represents the dual-domain enhanced feature map.

[0065] Figure 3 The dual-domain attention enhancement mechanism mentioned above is illustrated. Through the dual-domain attention enhancement process, not only the high-frequency detail features closely related to printing defects are highlighted from the perspective of the frequency domain, but also the important spatial information in the feature map is strengthened from the perspective of the spatial domain. The resulting dual-domain enhanced feature map can more accurately capture the diversity and complexity of printing defects, thereby providing a more solid foundation for subsequent defect detection.

[0066] Step S3: Perform multi-scale feature fusion processing on the enhanced feature maps of other layers except the first layer, and integrate context information to obtain a multi-scale context interaction enhanced feature map;

[0067] Combining the five-layer feature map example above, we can expand on this by performing multi-scale feature fusion on the second to fifth layers of feature maps, enhanced by dual-domain attention. Increasing the depth of feature extraction can lead to the degradation of key information in the original data, disrupting the temporal and spatial continuity of the data, and causing discontinuities in time series and geographic distribution, thus affecting the accuracy and reliability of subsequent analysis.

[0068] In a specific embodiment, as shown in formula (3), multi-scale feature fusion includes enhancing the multi-scale representation capability through a progressive cross-scale feature aggregation mechanism. First, the shallow features Perform single-branch convolution processing to generate ; then Resize to The scale of the layer is fused by double-branch convolution. get ; Then resize the output of the first two stages to The resolution of , together with the current layer features, is input into the three-branch convolution module to generate ;Finally adjust the feature size of all previous layers to the deepest layer The size of the network is fused with complete multi-scale information output through four-branch convolution. The whole process uses bilinear interpolation to unify the spatial dimension, uses 3×3 convolution for channel alignment, and forms a feature pyramid from fine-grained to coarse-grained by gradually superimposing shallow details and deep semantic features, achieving cross-level feature complementarity and information flow, and effectively improving the model's adaptability to complex scale changes. Figure 4 Schematic diagram of multi-scale feature fusion.

[0069]

[0070] in, represents the multi-scale feature map enhanced by dual domains, Represents the output feature map after multi-scale feature fusion, Indicates feature map size adjustment.

[0071] This strategy uses a cross-layer feature complementation mechanism to adaptively fuse high-resolution detail information from shallow networks (such as the edge texture of printed defects) with low-resolution semantic features from deep networks (such as the overall structural context), effectively alleviating the information degradation caused by progressive downsampling. This enables accurate restoration of image details during reconstruction, thereby improving defect segmentation accuracy.

[0072] Step S4: performing cross-scale refinement processing on the multi-scale context interaction enhanced feature map to obtain a refined feature map;

[0073] To expand on the previous example, the four layers of feature maps after multi-scale fusion can be divided into two groups: two shallow layers of feature maps in one group, and two deep layers of feature maps in another group, and then fed into the cross-scale wavelet refinement module. The core concept of this step is to decompose the feature maps into different frequency subbands using a cross-scale wavelet refinement network, and then combine it with multi-head cross-attention to achieve refined interaction of cross-scale features.

[0074] In order to solve the semantic gap and edge fuzzy problems in the decoding process, the feature map after multi-scale feature fusion is input into the cross-scale wavelet refinement module, combined with Figure 5 As shown in the figure, this module achieves cross-scale feature fusion through frequency domain feature decoupling and multi-head cross-attention mechanism. This module integrates wavelet transform into the attention calculation process, effectively decoupling low-frequency structural features from high-frequency texture features, significantly improving the integrity and edge consistency of multi-scale feature fusion.

[0075] In a specific embodiment, the cross-scale wavelet refinement module includes dividing the four-layer variation feature map into two groups: two low-layer feature maps as one group and two high-layer feature maps as another group, and inputting them into the cross-scale wavelet refinement module. Perform feature decoupling and downsampling. As shown in formula (4), in this process, is decomposed into four wavelet subbands: 、 、 and . The sub-band retains the low-frequency information of the structure of printing defects on the surface of cigarette packaging paper. and Sub-bands capture directional edge features, Subbands enhance the high frequency response of texture details. Four wavelet subbands are spliced ​​along the channel dimension. For details, please refer to Figure 6 Schematic example of discrete wavelet transform.

[0076]

[0077] As shown in formula (5), the discrete wavelet transform Transform to generate key-value pairs , and the query vector It is generated from the original variation feature map by exchanging the query vector between the feature maps of different scales. Then, the output feature map is processed with 1×1 convolution.

[0078]

[0079] in, represents the feature dimension, and Represent the outputs of two groups of multi-head cross attention, and Represent the refined outputs of the shallow and deep group feature maps, respectively.

[0080] Different from the layer-by-layer upsampling of traditional decoders, cross-scale information interaction is achieved through multi-head cross attention. At the same time, this module also realizes the deep coupling of spatial and frequency domain features. The direction-sensitive feature representation is enhanced by frequency sub-band reorganization. The attention mechanism introduces spatial adaptability in the feature reconstruction process, effectively suppressing the aliasing effect in the traditional inverse transform. In addition, when the discrete wavelet transform operation is performed on the feature map, the feature map is also downsampled, which reduces the computational cost of the model during calculation. The complexity of traditional self-attention is Resolution , the input dimension is compressed to , reducing the computational complexity of cross attention to .

[0081] Step S5: Concatenate and reconstruct the first-layer feature map enhanced by dual-domain attention and the refined feature map to obtain a reconstructed image;

[0082] To elaborate, in a specific implementation, as shown in formula (6), the two sets of output feature maps after wavelet refinement are first upsampled to match the spatial size of the first-layer feature map after dual-domain attention enhancement. Subsequently, the upsampled feature map and the first-layer feature map after dual-domain attention enhancement are spliced ​​in the channel dimension to fuse multi-dimensional information. Finally, the upsampling operation is performed again to further improve the spatial resolution and detail richness of the feature map, thereby obtaining the final reconstructed feature map. This process effectively integrates feature information from different sources, allowing the final feature map to more comprehensively reflect the details and semantic information of the image, providing a solid foundation for subsequent classification and localization tasks.

[0083]

[0084] in, represents the shallowest feature map after dual-domain attention enhancement, Represents the final reconstructed feature map.

[0085] Step S6: Using the reconstructed image, determine whether there are surface defects and the location of the defects in the original image of the cigarette packaging paper to be tested.

[0086] In this step, as shown in formula (7), the reconstructed feature map after cross-scale refinement is input into the detection head of the model. First, convolution, batch normalization, and nonlinear activation operations are performed through the feature enhancement layer to improve feature expression capabilities. A dropout layer is then used to randomly mask some neuron outputs to suppress the risk of model overfitting. Finally, a 1×1 convolutional layer is used to map feature channels and generate a pixel-level probability map. The probability map is binarized based on a preset threshold, and a defect segmentation mask is output to accurately locate the pixel-level outline of the printed defects on the cigarette packaging paper.

[0087]

[0088] in, is the mask matrix, is the drop probability, Represents the position in the output segmentation map The value at .

[0089] Combine Figure 7 A case study of cigarette packaging paper surface printing defect detection based on the solution of the present invention is shown. The visualization results show that the cigarette packaging paper surface defect detection method proposed in the present invention can accurately locate and identify cigarette packaging paper surface printing defects of various sizes and types, meeting the requirements of cigarette packaging paper surface printing defect detection in actual production scenarios.

[0090] In summary, the main design concept of the present invention is to extract a multi-scale feature map of the image of the cigarette packaging paper to be tested; use dual-domain attention to enhance the area where the surface defects are located in the multi-scale feature map; perform multi-scale feature fusion processing on the enhanced deeper feature map, and integrate context information to obtain a multi-scale context interaction enhanced feature map; perform cross-scale refinement processing on the multi-scale context interaction enhanced feature map to obtain a refined feature map; splice the enhanced first-layer feature map with the refined feature map to obtain a reconstructed image; and use the reconstructed image to determine whether there are surface defects and the defect locations in the image of the cigarette packaging paper to be tested. The present invention integrates frequency domain and spatial domain information to realize multi-scale information interaction, significantly improving the accuracy of cigarette packaging paper defect detection, especially improving the detection accuracy of cigarette packaging paper surface printing defects and the ability to locate printing defect edges, solving the problems of information loss, fuzzy details and semantic gaps in previous cigarette packaging paper surface defect detection schemes.

[0091] If the expressions expressing directions are mentioned in the embodiments of the present invention, they are relative concepts based on the embodiments. In addition, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a, b and c, where a, b, c can be single or multiple.

[0092] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings, but the above is only a preferred embodiment of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred modes can be reasonably combined and matched into a variety of equivalent schemes by those skilled in the art without departing from or changing the design ideas and technical effects of the present invention; therefore, the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.

Claims

1. A method for detecting surface defects of cigarette packaging paper based on dual-domain attention, characterized in that: include: Perform feature extraction on the original image of the cigarette packaging paper to be tested to obtain a multi-scale feature map; Using dual-domain attention to enhance the region of interest of the multi-scale feature map, the region of interest is the area where the surface defects of the cigarette packaging paper are located; The enhanced feature maps of other layers except the first layer are subjected to multi-scale feature fusion processing, and context information is integrated to obtain multi-scale context interaction enhanced feature maps; Perform cross-scale refinement processing on the multi-scale context interaction enhanced feature map to obtain a refined feature map; The first-layer feature map enhanced by dual-domain attention is concatenated with the refined feature map to obtain a reconstructed image. The reconstructed image is used to determine whether there are surface defects and the location of the defects in the original image of the cigarette packaging paper to be tested.

2. The cigarette packaging paper surface defect detection method based on dual-domain attention according to claim 1 is characterized in that: The methods for performing dual-domain attention enhancement on multi-scale feature maps include: The multi-scale feature map is input into the wavelet attention enhancement network, and the feature map is decoupled by discrete wavelet transform to generate the frequency domain enhanced feature map; The multi-scale feature map is input into the spatial attention enhancement network, and the spatial saliency weight is extracted through the global pooling operation to generate the spatially enhanced feature map; The frequency domain enhanced feature map and the spatial domain enhanced feature map are added pixel by pixel to obtain the dual-domain attention enhanced feature map.

3. The cigarette packaging paper surface defect detection method based on dual-domain attention according to claim 2, characterized in that: Generating the frequency domain enhanced feature map includes: The input multi-scale feature map is decoupled into four frequency sub-band feature maps through discrete wavelet transform, including low-frequency sub-band, horizontal high-frequency sub-band, vertical high-frequency sub-band and diagonal high-frequency sub-band; The horizontal, vertical and diagonal high-frequency sub-bands are added element by element, and the frequency domain attention map is generated through the Sigmoid activation function; Multiplying the frequency domain attention map by the low-frequency subband feature map pixel by pixel to obtain a wavelet enhanced feature map; The wavelet enhanced feature map is concatenated with the original input multi-scale feature map along the channel dimension to obtain a frequency domain enhanced feature map.

4. The cigarette packaging paper surface defect detection method based on dual-domain attention according to claim 2, characterized in that: Generating the spatially enhanced feature map includes: Perform global average pooling and global maximum pooling operations on the input multi-scale feature map to obtain two pooled feature maps; The two pooled feature maps are concatenated along the channel dimension, and a spatial attention weight map is generated through a convolutional layer and a Sigmoid activation function; The spatial attention weight map is multiplied pixel by pixel with the multi-scale feature map of the original input and the feature response strength is adjusted to obtain a spatially enhanced feature map.

5. The cigarette packaging paper surface defect detection method based on dual-domain attention according to claim 1, characterized in that: The obtaining of a multi-scale context interaction enhanced feature map comprises: The deep feature map enhanced by dual-domain attention is used to interact with contextual information through a progressive cross-scale feature aggregation mechanism; Perform single-branch convolution on shallow features to generate preliminary fusion features; After resizing the initial fusion features, they are fused with the next layer features through a two-branch convolution; After the previous fusion features are scaled and matched with the current layer features, they are subjected to multi-branch convolution processing with the current layer features and aggregated layer by layer; Generate multi-scale contextual interaction enhanced feature maps for adaptive fusion of shallow details and deep semantics.

6. The cigarette packaging paper surface defect detection method based on dual-domain attention according to claim 1, characterized in that: Obtaining the refined feature map includes: Divide the multi-scale context interaction enhancement feature map into shallow group feature maps and deep group feature maps; Each set of feature maps is fed into a cross-scale wavelet-refined cross-attention network; The discrete wavelet transform is used to perform feature decoupling and lossless downsampling on each group of feature maps to obtain multiple decomposed wavelet sub-bands. The multi-head cross-attention mechanism is combined to interact information between feature maps of different scales and output the refined feature maps.

7. The cigarette packaging paper surface defect detection method based on dual-domain attention according to claim 1, characterized in that: Obtaining the reconstructed image includes: The two sets of feature maps after cross-scale wavelet refinement are upsampled separately to match the size of the first layer feature map after dual-domain attention enhancement; The upsampled refined feature map is concatenated with the first-layer attention-enhanced feature map in the channel dimension to form an integrated feature map. The integrated feature map is upsampled again to restore it to its original size, and finally the reconstructed feature map is obtained.

8. The method for detecting surface defects of cigarette packaging paper based on dual-domain attention according to any one of claims 1 to 7, characterized in that: The feature extraction includes: using the backbone network of a pre-built cigarette packaging paper detection model to extract features from the original image, and obtaining a multi-scale feature map with decreasing spatial size and increasing channel number from the first layer to the last layer.

Citation Information

Cited By

  • Intelligent trackside detection method based on unsupervised reconstruction normal form

    CN121120660A