A plastic film surface defect detection method based on multi-scale feature fusion

CN122841301APending Publication Date: 2026-09-29HUBEI JINDE PACKAGING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610974318.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]为了解决现有技术存在的背景干扰强、缺陷对比度低、微小缺陷细节易丢失、多尺度缺陷分割精度不足的问题,本发明提供一种基于多尺度特征融合的塑料薄膜表面缺陷检测方法,该方法包括:

Benefits of technology

本发明针对塑料薄膜高反光、半透明的物理特性,通过将背景噪声从缺陷信号中分离并定向压制,消除了由于光照不均、材质底纹造成的误检问题,具备较强的工业级抗背景干扰能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841301A_ABST
    Figure CN122841301A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of defect detection, and particularly relates to a plastic film surface defect detection method based on multi-scale feature fusion. The method comprises: performing dual tree complex wavelet transform on a plastic film image to obtain global background prior and multi-channel detail features; constructing a double-flow backbone network, setting a multi-hole rate parallel receptive field module, a first flow extracting global context features of an original image, and a second flow extracting detail structure features; through a top-down fusion path, using a gate unit to generate temporary spatial weights and channel weights, combining background suppression attention maps to form spatial weights, and performing spatial screening and channel weighting on low-level features; inputting the preliminary fused features into a source perception residual correction module, combining double-flow same-level source features and preliminary fused features to calculate compensation residuals, obtaining a feature map set and performing classification prediction, and outputting a defect detection result. The present application improves defect boundary integrity and stability of pixel-level segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of defect detection technology, and specifically to a method for detecting surface defects in plastic films based on multi-scale feature fusion. Background Technology

[0002] In the actual production and manufacturing process of plastic film, due to various factors such as the precision of processing equipment, fluctuations in process parameters, impurities in raw materials, and changes in the external environment, the film surface is prone to defects or flaws such as scratches, crystal points, holes, bubbles, and foreign matter spots. Traditional manual visual inspection methods suffer from problems such as high labor intensity, low inspection efficiency, and susceptibility to personnel fatigue and subjective experience. In contrast, machine vision-based defect detection methods have the advantages of objectivity and automation. However, in actual industrial scenarios, plastic films usually have characteristics such as high reflectivity, transparency, or translucency. The image acquisition process is easily affected by non-uniform lighting, equipment reflection, and the perspective of the underlying background, resulting in problems such as complex backgrounds, high noise, and weak contrast between defects and the background in the film images.

[0003] Meanwhile, plastic film surface defects are diverse in type, irregular in shape, and span a wide range of scales, potentially manifesting as large-area contour anomalies or minute dot-like or linear flaws. Under complex backgrounds and varying noise conditions, existing models are easily affected by irrelevant background interference, making it difficult to stably focus on the actual defect areas. During multi-scale feature transfer and deep abstraction, the spatial structure and edge details of minute defects may be smoothed or weakened, leading to missed detections or inaccurate boundary segmentation. Therefore, a plastic film surface defect detection scheme is needed that can suppress background interference at a global level, integrate multi-scale spatial features and detailed features, and compensate for the detailed information lost during transfer, in order to improve the accuracy of defect detection and the stability of segmentation results. Summary of the Invention

[0004] To address the problems of strong background interference, low defect contrast, easy loss of details in minute defects, and insufficient multi-scale defect segmentation accuracy in existing technologies, this invention provides a method for detecting surface defects in plastic films based on multi-scale feature fusion. This method includes:

[0005] S1. Acquire the plastic film image and perform dual-tree complex wavelet transform to decompose the low-frequency sub-band as the global background prior, and integrate the high-frequency directional sub-band into multi-channel detail features; S2. Construct a dual-stream backbone network, in which each stage of the dual-stream backbone network has a built-in parallel receptive field module composed of multiple convolutional branches with different dilation rates; S3. Establish a top-down fusion path, generate temporary spatial weights and channel weights matching the number of low-level feature channels to be fused, combine with the global background prior to generate an initial feature map, and simultaneously construct a background suppression attention map, combine with the temporary spatial weights to obtain spatial weights, and then obtain preliminary fusion features; S4. Based on the preliminary fusion features, calculate the compensation residual, and then restore the original details to obtain a feature map set, perform pixel-level classification prediction on the feature map set, and output the plastic film surface defect segmentation result.

[0006] This invention separates the background illumination and defect details of the thin film through dual-tree complex wavelet transform at the front end. Combined with a dual-stream backbone network, it achieves independent extraction of global semantic and microstructural features. In the feature fusion stage, by introducing a background suppression attention map, it effectively filters background interference caused by highly reflective and transparent materials. At the same time, with the residual correction mechanism, it recovers lost micro-defect features before output. This effectively solves the problems of strong background interference, easy omission of low-contrast defects, and blurred boundaries caused by downsampling in traditional visual inspection, and improves the accuracy of pixel-level segmentation of thin film defects.

[0007] Further, the acquisition of the plastic film image and the performance of dual-tree complex wavelet transform specifically includes: inputting the single-channel brightness image corresponding to the plastic film image into a preset dual-tree separation analysis branch; performing one-dimensional low-pass and high-pass filtering and downsampling operations sequentially along the horizontal and vertical directions of the image at each decomposition level; after the decomposition loop, extracting the top-level low-pass output as the low-frequency sub-band; converting the low-frequency sub-band into a real-valued low-frequency amplitude map as the global background prior; and performing real-valued processing on the complex high-frequency output coefficients generated at each level to obtain a high-frequency amplitude feature map; upsampling the high-frequency amplitude feature map to the spatial resolution of the plastic film image; and stitching it together according to the channel dimension to obtain the multi-channel detail features.

[0008] This invention utilizes a pre-defined dual-tree separation analysis branch for multi-level filtering. While ensuring translation invariance in the decomposition process, it effectively extracts high-frequency amplitude features in various directions, providing high-quality prior information on edge mutations for subsequent neural networks. This reduces the difficulty for deep learning models to fit complex background noise from scratch and improves the purity of front-end feature extraction.

[0009] Furthermore, the construction of the dual-stream backbone network includes: in the first stream, retaining the initial large convolutional kernels and max-pooling downsampling operations of the pre-trained convolutional neural network to establish a main feature path for extracting global contextual features; in the second stream, adopting a convolutional layer architecture of the same depth as the first stream, with the input channel receiving the multi-channel detailed features, and performing downsampling operations synchronously in both streams, maintaining the same spatial resolution at each corresponding feature level, and obtaining structural features through parallel extraction.

[0010] Furthermore, each stage of the dual-stream backbone network incorporates a parallel receptive field module consisting of multiple convolutional branches with different dilation rates. This includes setting up three parallel convolutional branches at the same level node for feature extraction. Each branch uses a convolutional kernel of the same size to receive the same input feature map. The multi-scale feature maps extracted by the three parallel branches are spliced ​​and fused in the channel dimension, and the output is used as the feature representation of the corresponding level.

[0011] This invention expands the receptive field of feature extraction without increasing the number of additional model parameters by introducing multiple convolutional branches with different hole rates in parallel at the same level of feature extraction. This allows the model to observe the film surface simultaneously with different perspectives, covering both large-area hole outlines and fine pinhole scratches, effectively improving the model's adaptability to defects with large spans and multiple scales.

[0012] Further, the generation of temporary spatial weights and channel weights matching the number of low-level feature channels to be fused includes: performing global average pooling on the input high-level features to compress the spatial dimension, obtaining global channel statistics, and then sequentially passing the channel weights through a multilayer perceptron with an output dimension equal to the number of low-level feature channels to be fused and an activation function; inputting the high-level features into a two-dimensional convolutional layer, and after processing by the activation function, generating a two-dimensional matrix as the temporary spatial weights.

[0013] Furthermore, the construction of the background suppression attention map includes: using a bilinear interpolation algorithm to proportionally enlarge the low-frequency sub-band, which serves as the global background prior, according to the length and width dimensions of the low-level features to be fused, to obtain an enlarged feature map; inputting the enlarged feature map into a sub-network composed of convolutional layers and activation functions, outputting the initial feature map, and subtracting the initial feature map from it using a unit tensor of the same size as the initial feature map to obtain the background suppression attention map, which is used to suppress the feature response values ​​of the background noise region.

[0014] This invention reverses the low-frequency signal representing the normal material response into a background suppression mask, which can suppress the feature activation values ​​of non-defect areas in the spatial dimension, avoiding interference from complex backgrounds such as equipment reflections and transparent patterns on the detection results.

[0015] Further, the process of restoring the original details and obtaining the feature map set includes: extracting the source features corresponding to the same level, which are output by the first stream before fusion, and the source features output by the second stream; concatenating the two sets of source features with the preliminary fusion features in the channel dimension to obtain composite features; inputting the composite features into a correction module for dimensionality reduction convolution and feature extraction convolution to obtain the compensation residual used to compensate for spatial and semantic information loss; and adding the compensation residual to the preliminary fusion features element-wise to complete the feature residual correction and restoration, thereby obtaining the feature map set.

[0016] This invention accurately repairs high-frequency contour wear caused by gating or multi-layer downsampling by calling the original features output from the first and second streams of the network front-end, splicing them with the fused features and adding the residuals, thus ensuring the effectiveness of edge segmentation for minute defects such as crystal points and fine hairs.

[0017] Further, obtaining the preliminary fusion features includes: upsampling the temporary spatial weights to the same spatial resolution as the low-level features to be fused, and multiplying them element-wise with the background suppression attention map to obtain composite spatial weights; using the composite spatial weights to perform spatial filtering on the low-level features to be fused, and using the channel weights to perform channel weighting on the low-level features to be fused to obtain the preliminary fusion features.

[0018] Furthermore, pixel-level classification prediction of the feature map set includes: uniformly upsampling feature maps of different resolutions in the feature map set to a preset decoding resolution, mapping them to the same number of channels, and concatenating the aligned multi-level features along the channel dimension to form a comprehensive segmentation feature map; inputting the comprehensive segmentation feature map into multiple convolutional layers and activation layers in sequence for feature decoding to generate a classification score map, selecting the category with the highest probability in the channel dimension as the category label of the corresponding pixel, and outputting the segmentation result of the surface defect of the plastic film.

[0019] Furthermore, the method for detecting surface defects of plastic films based on multi-scale feature fusion further includes: acquiring plastic film defect images with pixel-level labeled masks as training samples, inputting the original images and the multi-channel detail features into a dual-stream backbone network respectively to obtain classification score maps of the corresponding samples; comparing the classification score maps with the pixel-level defect labels of the training samples, constructing a joint loss function composed of cross-entropy loss and Dice loss weighted together, and performing end-to-end parameter updates on the network model based on the joint loss function.

[0020] The present invention has the following technical effects: This invention addresses the high reflectivity and semi-transparency of plastic films by separating and directionally suppressing background noise from defect signals, thus eliminating false detections caused by uneven lighting and material texture, and possessing strong industrial-grade anti-background interference capabilities. This invention achieves complete capture of defects across a wide range and at multiple scales, from macroscopic morphology to microscopic texture. Through the dynamic field of view of parallel dilated convolution and the synchronous downsampling of the dual-stream structure, it ensures that defects of different sizes can find the most matching receptive field in the network, effectively reducing the false negative rate of minor defects. Attached Figure Description

[0021] Figure 1 This is a flowchart of a method for detecting surface defects in plastic films based on multi-scale feature fusion, provided by an embodiment of the present invention. Figure 2 This is a diagram showing the resolution variation of a dual-stream network provided in an embodiment of the present invention; Figure 3 This is a comparison chart of ablation experiments provided in the embodiments of the present invention. Detailed Implementation

[0022] This invention provides a method for detecting surface defects in plastic films based on multi-scale feature fusion, referring to... Figure 1 This includes steps S1-S4: S1: Data Acquisition and Preprocessing.

[0023] Specifically, the plastic film image is acquired and subjected to dual-tree complex wavelet transform to decompose the low-frequency sub-band as the global background prior, and the high-frequency directional sub-band is integrated into multi-channel detail features.

[0024] First, a digital image of the plastic film surface is acquired using an industrial linear scan camera. This digital image is read into a three-channel color image matrix data and then converted into a single-channel brightness image for performing a dual-tree complex wavelet transform. The three-channel color image matrix data is retained as the input to the first-order network. Next, a four-level dual-tree complex wavelet transform decomposition is performed on the single-channel brightness image. In this embodiment, the preferred wavelet filter bank is a dual-tree complex wavelet filter bank composed of a near-symmetric filter and a Q-shift filter. The low-frequency approximation coefficient matrix of the fourth level is extracted from the decomposition result. When the low-frequency approximation coefficient matrix contains complex coefficients, the sum of the squares of the real part and the square root of the imaginary part is calculated to obtain the real-valued low-frequency amplitude map, which serves as the global background prior.

[0025] Subsequently, high-frequency detail coefficient matrices in six different directions generated by each level of decomposition are extracted. For each direction of high-frequency detail coefficient matrix, the sum of squares of the elements of the real and imaginary part matrices of the complex coefficients is calculated, and the square root algorithm is used to obtain the corresponding amplitude feature map. After upsampling the amplitude feature maps of each direction at different levels to a uniform spatial resolution, they are concatenated and stitched along the channel dimension to generate multi-channel detail features.

[0026] In some implementations, the plastic film image is acquired and subjected to dual-tree complex wavelet transform to decompose it into a low-frequency sub-band as a global background prior. The high-frequency directional sub-bands are then integrated into multi-channel detail features. This process includes: inputting the single-channel brightness image corresponding to the plastic film image into a dual-tree separation analysis branch containing a first real-valued filter bank and a second real-valued filter bank; performing one-dimensional low-pass and high-pass filtering and downsampling operations sequentially along the horizontal and vertical directions of the image at each decomposition level; after four consecutive decomposition cycles, extracting the top-level low-pass output as the low-frequency sub-band; converting the low-frequency sub-band into a real-valued low-frequency amplitude map as a global background prior; and performing real-valued processing on the complex high-frequency output coefficients generated at each level in six directions (±15°, ±45°, ±75°) to obtain high-frequency amplitude feature maps in each direction. After upsampling the high-frequency amplitude feature maps in each direction to the spatial resolution of the plastic film image, they are stitched together according to the channel dimension to form multi-channel detail features.

[0027] When performing feature decomposition on images of plastic films with surface defects such as scratches, bubbles, or holes, such as images with a spatial resolution of 1024×1024 pixels, this embodiment employs a dual tree complex wavelet transform with approximate shift invariance. Specifically, two different filter banks that meet the complete reconstruction conditions are pre-configured. In this embodiment, a Q-shift filter is preferred, which constitutes TreeA branch and TreeB branch respectively. The outputs of the two branches are combined to form the real part approximation and imaginary part approximation of the complex wavelet coefficients.

[0028] At each decomposition level, the two branches first perform one-dimensional filtering on the input signal in the horizontal direction, including low-pass and high-pass filtering, and then downsample in this direction at a ratio of 2:1. The same one-dimensional filtering and downsampling operation is repeated in the vertical direction. Taking an input image with a resolution of 1024×1024 as an example, after the first decomposition, the output resolution is reduced to 512×512. This process is repeated four times. When the fourth decomposition level is reached, a complex low-frequency component with a resolution of 64×64 is generated. After amplitude processing of the complex low-frequency component, a real-valued low-frequency sub-band representing the overall illumination and material distribution of the thin film is obtained, and this is used as the global background prior.

[0029] Simultaneously, in the first to fourth decomposition levels, the coefficients output by TreeA and TreeB branches at the high-pass filter positions are extracted respectively, and the high-pass coefficients of the two branches at the same level and the same filter combination position are combined to form complex high-frequency wavelet coefficients. Using the directional selectivity of the two-dimensional dual-tree complex wavelet filter bank, high-frequency components in six directions (+15°, -15°, +45°, -45°, +75°, -75°) are obtained. Subsequently, a bicubic interpolation algorithm is used to uniformly upsample the 24 high-frequency feature maps obtained from the first layer (512×512), the second layer (256×256), the third layer (128×128), and the fourth layer (64×64) to a spatial resolution of 1024×1024.

[0030] Finally, the feature maps are stitched together along the channel dimension to obtain a tensor with a size of 1024×1024 and a depth of 24 channels, which serves as a multi-channel detail feature that preserves the edge contours of defects and local high-frequency abrupt changes at multiple scales.

[0031] S2: Construction of a dual-stream backbone network.

[0032] Specifically, a dual-stream backbone network is constructed, in which each stage of the dual-stream backbone network has a built-in parallel receptive field module consisting of multiple convolutional branches with different dilation rates.

[0033] First, two parallel ResNet50 network architectures are constructed to form a dual-stream backbone network. The plastic film image is normalized and then input into the first-stream network. Initial downsampling is completed through the initial convolutional layer and the max pooling layer. Then, multi-scale global context features are extracted sequentially through convolutional operations in four residual network stages.

[0034] Multi-channel detail features are input into the second-stream network. Since the multi-channel detail features have been uniformly upsampled to the same spatial resolution as the original plastic film image, the second stream adopts the residual network stage corresponding to the first stream and performs downsampling operation synchronously to maintain the same spatial resolution at the corresponding level. The parallel receptive field module is added to the output of each residual network stage of the dual-stream backbone network to enhance the receptive field of the corresponding level features at multiple scales.

[0035] The parallel receptive field module creates three parallel convolutional branches. The first branch is set to a regular 3×3 convolutional kernel with a dilation rate of 1, and the second and third branches are set to 3×3 dilated convolutional kernels with dilation rates of 3 and 5, respectively. The feature maps output by each branch are sequentially input into the batch normalization layer and the ReLU activation layer for processing, then concatenated along the channel dimension, and mapped to the corresponding level's output feature map through a 1×1 convolutional layer.

[0036] In some implementations, constructing a dual-stream backbone network involves: in the first stream, retaining the initial 7×7 large convolutional kernels and max-pooling downsampling operations of the pre-trained convolutional neural network to establish a main feature path for extracting global contextual features; in the second stream, adopting a convolutional layer architecture of the same depth as the first stream, with the input channels receiving multi-channel detailed features, and performing downsampling operations synchronously in both streams to maintain the same spatial resolution at each corresponding feature level, and obtaining structural features through parallel extraction.

[0037] The dual-stream backbone network consists of a first-stream network architecture initialized based on a pre-trained convolutional neural network, and a second-stream network architecture after adjusting the input layer according to 24-channel multi-channel detail features. The input of the first-stream network architecture is a 3-channel original RGB plastic film image, and the input of the second-stream network architecture is a 24-channel multi-channel detail feature. Both the first-stream and second-stream network architectures contain residual stages and convolutional downsampling layers of equal depth. The initial stage of the first-stream network contains a 7×7 large-size convolutional kernel and a max-pooling layer. The initial stage of the second-stream network uses the architecture corresponding to the first-stream network for parallel processing, and adjusts the number of input channels of the first convolutional layer to 24 channels. The output of the dual-stream backbone network is a feature map that maintains the same spatial resolution at each corresponding feature level and contains macroscopic morphological features and fine structural features respectively.

[0038] This dual-stream backbone network uses ResNet50 as its basic architecture and extends it to two branches. The first stream processes the original RGB plastic film image with an input layer of 3 channels. At the beginning of this branch, a standard 7×7 large-size convolutional kernel is retained with a stride of 2 and an output channel count of 64. Combined with a 3×3 max pooling layer with a stride of 2, downsampling is achieved to build a larger receptive field, resulting in a main feature pathway focused on acquiring macroscopic morphology and contextual semantic features.

[0039] In this process, the spatial resolution of a 1024×1024 input image is reduced to 256×256 after passing through the initial layer. Meanwhile, the second stream serves as a detail feature extraction branch, and the network depth and the distribution structure of the four residual stages are consistent with the first stream. The input layer is configured with 24 channels, receiving a 24-channel multi-channel detail feature tensor generated by the dual-tree complex wavelet transform.

[0040] The second stream employs a synchronous control strategy during convolution and stride downsampling to ensure that the output resolution of each corresponding feature layer is the same. Except for adjusting the number of input channels to 24, the stride and max pooling settings of the first layer convolution of the second stream are consistent with those of the first stream, so that the two streams output the same spatial resolution at the corresponding stages. For example, after stages 2, 3, and 4 of the first and second streams, the feature map spatial sizes output by the two branches remain at 256×256, 128×128, and 64×64, respectively. Figure 2This is a resolution variation diagram of the dual-stream network provided in the embodiment of the present invention. It can be seen that this parallel setting reduces the mutual influence between the background smoothing information of the thin film and the high-frequency boundary information of the defects, while ensuring the acquisition of structures of different abstract dimensions.

[0041] In some implementations, each stage of the dual-stream backbone network incorporates a parallel receptive field module consisting of multiple convolutional branches with different dilation rates. This includes: setting three parallel convolutional branches at the same level node for feature extraction, with each branch using a 3×3 convolutional kernel to receive the same input feature map; the first branch has a dilation rate of 1 for standard convolution, while the second and third branches have dilation rates greater than 1 that increase progressively, thus multiplying the receptive field range without increasing the number of parameters; and concatenating and fusing the multi-scale feature maps extracted by the three parallel branches along the channel dimension, outputting the feature representation for the corresponding level.

[0042] The input to the parallel receptive field module is a high-order feature map output by a dual-stream backbone network. The structure includes three parallel two-dimensional convolutional branches with different dilation rates, as well as subsequent feature concatenation layers and 1×1 dimensionality reduction convolutional layers. The output is a corresponding layer feature representation with differentiated receptive fields.

[0043] The parallel receptive field module is embedded into the output of the feature extraction stage of the dual-stream backbone network. In this embodiment, it is preferred to embed it into the output of the mid-to-high-order feature stage to enhance the model's ability to detect the morphology of thin film defects of different sizes. If the current input feature map size of the module is 64×64 and the number of channels is 1024, the tensor is synchronously split into three parallel two-dimensional convolution branches.

[0044] The first branch is configured as a standard 3×3 convolution with a dilation rate of 1, responsible for extracting local neighboring pixel association information, with an actual receptive field of 3×3; the second branch sets the dilation rate to 3, and fills the weights with two zero-value matrices during actual calculation, thereby expanding its receptive field to 7×7 to cover medium-sized material textures while keeping the nine learning parameters unchanged; the third branch increases the dilation rate to 5, expanding the receptive field to 11×11, covering larger-scale global deformation and defect trends.

[0045] It should be noted that during the operation, the number of output channels for each of the three branches can be uniformly set to 256 channels. After the operation of each branch is completed, it is processed by batch normalization layer and ReLU activation function. The three sets of feature maps with differential receptive fields are spliced ​​along the channel dimension, and a 1×1 standard convolution kernel is added to remap them back to the original 1024 dimensions. As a result, the corresponding layer feature representation output covers microscopic details and improves the network's adaptability to defects of different scales.

[0046] S3: Preliminary fusion characteristics determined.

[0047] Specifically, a top-down fusion path is established, temporary spatial weights and channel weights matching the number of low-level feature channels to be fused are generated, an initial feature map is generated by combining global background priors, a background suppression attention map is constructed, and spatial weights are obtained by combining temporary spatial weights, thus obtaining preliminary fused features.

[0048] First, before entering the top-down fusion path, the feature maps output by the highest layer of the first-level network and the feature maps output by the second-level network are concatenated along the channel dimension and fused through a 1×1 convolutional layer, serving as the high-level feature input for the top-down path. Simultaneously, the features output by the first-level network at the corresponding resolution level are used as low-level features to be fused, participating in weight selection via lateral connections. The corresponding level features output by the second-level network are retained as the source features for subsequent residual correction. This embodiment employs a Feature Pyramid Network (FPN) architecture to construct a top-down fusion path. First, the high-level features output by the backbone network are reduced in dimensionality using 1×1 convolutions to match the number of channels with the layers to be fused. Next, the dimensionality-reduced high-level features are input into the channel weight generation branch and the spatial weight generation branch, respectively. In the channel weight generation branch, global average pooling is performed on the dimensionality-reduced high-level features to obtain a global channel statistical vector, which is then input into a multilayer perceptron and a sigmoid activation layer to generate channel weights with values ​​ranging from 0 to 1. In the spatial weight generation branch, the dimensionality-reduced high-level features are input into a two-dimensional convolutional layer and processed by a sigmoid activation layer to generate temporary spatial weights with values ​​ranging from 0 to 1.

[0049] Simultaneously, the global background prior is upsampled using a bilinear interpolation algorithm to make its spatial size consistent with the low-level features to be fused. The upsampled global background prior is then input into a feature mapping subnetwork consisting of one or two convolutional layers, outputting a single-channel initial feature map. A unit tensor with the same size as the initial feature map and all elements equal to 1 is constructed, and the initial feature map is subtracted from the unit tensor to obtain a background suppression attention map. Subsequently, the temporary spatial weights are upsampled to the same spatial resolution as the low-level features to be fused passed through lateral connections in the bottom-up path, and multiplied element-wise with the background suppression attention map to obtain composite spatial weights. These composite spatial weights combine the localization effect of high-level semantic features on suspected defect regions and the suppression effect of low-frequency background priors on normal background regions to enhance defect response and reduce non-defect background noise.

[0050] Next, the composite spatial weights are multiplied element by element with the low-level features to be fused to complete the spatial dimension filtering. Then, the channel weights are applied along the channel dimension to the low-level features after spatial filtering using the tensor broadcasting mechanism to complete the channel weighting and obtain the preliminary fused features.

[0051] In some implementations, a gating unit generates temporary spatial weights and channel weights matching the number of channels of the low-level features to be fused based on high-level features. This includes: performing global average pooling on the input high-level features to compress the spatial dimension to 1×1, obtaining global channel statistics, and then sequentially passing the output dimension of a multilayer perceptron equal to the number of channels of the low-level features to be fused, followed by a sigmoid activation function, to generate the channel weights with values ​​ranging from 0 to 1. Simultaneously, the high-level features are input into a two-dimensional convolutional layer, processed by a sigmoid activation function, to generate a two-dimensional matrix corresponding to each pixel position with values ​​ranging from 0 to 1, which serves as the temporary spatial weights.

[0052] The input to this multilayer perceptron is one-dimensional global channel statistics of high-level features after global average pooling. The structure includes two fully connected layers for non-linear feature crossing and dimensionality reprojection. The output is a dimensionality-reduced feature vector that matches the number of low-level features to be fused.

[0053] In the top-down information fusion path, the gating unit generates weights through two parallel branches to regulate the spatial and channel responses of low-level features; if the dimension of the input deep semantic feature tensor is 64×64×1024, then the dimension of the low-level features to be fused is 128×128×512.

[0054] In the channel weight generation branch, global average pooling is first performed on the entire spatial plane of the high-level features to compress and generate a 1×1×1024 vector, which is used to detect the importance distribution of each semantic channel dimension. Then, the vector is input into a multilayer perceptron containing two fully connected layers. The first layer reduces the feature dimension to 64 dimensions and activates non-linear feature crossing. The second layer reprojects the dimension and outputs a 512-dimensional vector matching the number of low-level features. The Sigmoid function is used to generate 512-channel normalized weights with values ​​in the range of 0 to 1.

[0055] In the temporary spatial weight calculation branch, the gating mechanism focuses on the information of pixel position distribution. It takes 64×64×1024 high-level features as input and outputs a single-layer two-dimensional convolution with 1 channel. It uses 1×1 or 3×3 convolution kernels to extract the abnormal response distribution of different spatial coordinate points across all channels. Then, it uses the Sigmoid function to map the distribution to the range of 0 to 1, thereby generating a two-dimensional feature matrix with a single channel size of 64×64×1, which is defined as a temporary spatial weight map. In this weight map, pixel values ​​close to 1 correspond to suspected defect areas or abnormal response areas on the plastic film, which are used to guide the mask weighting of the high-resolution low-level map.

[0056] In some implementations, the global background prior is upsampled and an initial feature map is generated through a sub-network. A background suppression attention map is then constructed. This involves: using a bilinear interpolation algorithm to scale up the low-frequency sub-bands, which serve as the global background prior, according to the length and width of the low-level features to be fused, to obtain a scaled-up feature map; inputting the scaled-up feature map into a sub-network consisting of a 1×1 convolutional layer and a sigmoid activation function, outputting an initial feature map mapped to the 0-1 interval, and subtracting the initial feature map from a unit tensor with all elements set to 1 and the same size as the initial feature map to obtain the background suppression attention map, which is used to suppress the feature response values ​​of the background noise region through subsequent multiplication operations.

[0057] The input to this sub-network is a global background prior feature map that has been scaled up proportionally to the spatial scale. The structure includes a 1×1 point-directed convolutional layer without bias and a cascaded Sigmoid activation function layer. The output is an initial feature map that has been numerically normalized to the range of 0 to 1.

[0058] Using the low-frequency sub-band matrix obtained from the fourth layer of wavelet decomposition, such as a single channel with a resolution size of 64×64, background prior constraints are provided. The mean of grid coordinate weights is calculated using a bilinear interpolation algorithm. The low-frequency matrix is ​​then enlarged by a factor of 4 on both the width and height axes of the target layer to be fused, such as an output size of 128×128, to improve the spatial resolution to 256×256.

[0059] A magnified feature map, resized to 256×256×1, is input into a sub-network. This sub-network uses a 1×1 point-to-point convolutional layer with no bias to perform a smooth transformation on its feature amplitude distribution, and then normalizes it to the range of 0 to 1 using a Sigmoid activation function. This generates an initial feature map representing the response to normal lighting and material background. Values ​​closer to 1 indicate a more normal background response, while values ​​closer to 0 indicate a weaker background response or the possibility of an abnormal response. The parameters of this feature mapping sub-network are updated end-to-end with the pixel-level classification prediction loss, so that its output progressively represents the distribution of normal background response. Next, a full tensor with a value of 1 is constructed and subtracted pixel-by-pixel from the initial feature map to obtain a background suppression attention map. This map converts strong normal background values ​​into lower weight coefficients close to 0. These lower weight coefficients are used as a mask in the subsequent Hadamard multiplication calculation to suppress invalid low-frequency signals in the background region and improve the discriminative power of the fused features for the response to details of defects on the real plastic surface.

[0060] S4: Defect detection results are generated.

[0061] Specifically, based on the initial fusion features, the compensation residual is calculated to restore the original details and obtain the feature map set. The feature map set is then used for pixel-level classification prediction to output the plastic film surface defect segmentation results.

[0062] First, within the source-aware residual correction module, the low-level features to be fused from the first-stream output, the source feature matrix from the second-stream output, and the preliminary fusion features are combined along the channel dimension to form a composite correction feature. This composite correction feature is then fed into a compensation residual estimation network consisting of a 1×1 dimensionality reduction convolutional layer, a 3×3 feature extraction convolutional layer, a batch normalization layer, and a nonlinear activation layer to obtain feature residual values ​​used to compensate for fine edge information and spatial structure information, which serve as compensation residuals. Next, the compensation residuals are superimposed onto the preliminary fusion features through a matrix tensor bitwise addition operation, thereby completing the correction and repair of high-frequency structural details lost due to downsampling and the fusion process. After processing each multi-scale feature layer in sequence, a refined feature map is obtained.

[0063] Subsequently, refined feature maps containing multiple resolutions are input into the classification prediction module. This module employs a fully convolutional network structure, including a scale alignment unit, a channel integration unit, a feature decoding unit, and a classification output unit. First, the refined feature maps at different resolutions are uniformly upsampled to a preset decoding resolution using a bilinear interpolation algorithm. This preset decoding resolution can be set to the original plastic film image resolution or its proportionally scaled resolution. Each level of feature is then mapped to the same number of channels using a 1×1 convolutional layer. Next, the aligned multi-level features are concatenated along the channel dimension to form a comprehensive segmentation feature map. This comprehensive segmentation feature map is then sequentially input into two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation layer for local spatial relationship modeling and feature decoding. Finally, a classification score map is generated using a 1×1 convolutional layer. The number of channels in this classification score map is equal to the total number of background categories and plastic film surface defect categories.

[0064] During the training phase, this embodiment uses plastic film defect images with pixel-level labeled masks as training samples. The original RGB image and the multi-channel detail features obtained by dual-tree complex wavelet transform are input into the dual-stream backbone network. After feature fusion, residual correction, and classification prediction, a pixel-level classification score map is obtained. The classification score map is compared with manually labeled pixel-level defect labels, and a joint loss function composed of cross-entropy loss and Dice loss is constructed. The cross-entropy loss is used to constrain the prediction results of each pixel category, and the Dice loss is used to reduce the impact of uneven defect area on training. Then, the AdamW optimizer is used to update the end-to-end parameters of the dual-stream backbone network, gating unit, background suppression attention module, source perception residual correction module, and classification prediction module. The model parameters are selected by the mean intersection-union ratio and F1 score of the validation set.

[0065] During the inference phase, the image of the plastic film to be detected is input into the trained model. The classification and prediction module outputs the probability distribution of each pixel belonging to the background and various types of defects. The category with the highest probability is selected as the category label of the corresponding pixel in the channel dimension, generating a single-channel two-dimensional label matrix. This two-dimensional label matrix is ​​used as the result of plastic film surface defect segmentation.

[0066] In some implementations, the source features output by the first stream and the source features output by the second stream, along with the preliminary fused features, are concatenated at the same level before fusion. A compensation residual is calculated and added to the preliminary fused features to restore the original details and obtain a refined feature map. This includes: extracting the source features output by the first stream and the source features output by the second stream at the same level before fusion; concatenating the two sets of source features with the preliminary fused features along the channel dimension; inputting the concatenated composite features into a correction module containing 1×1 dimensionality reduction convolution and 3×3 feature extraction convolution to obtain the compensation residual used to compensate for spatial and semantic information loss; and adding the compensation residual to the preliminary fused features element-wise to complete the feature residual correction and restoration.

[0067] The input to the correction module is a stacked tensor formed by concatenating and recombining the first source feature, the second source feature, and the preliminary fused features in the channel dimension. The structure includes a 1×1 bottleneck convolutional layer configured in series for feature interaction dimensionality reduction, and a 3×3 standard spatial convolutional layer for extracting spatial topological relationships. The output is the dimension-matched compensation residual.

[0068] After top-down fusion and filtering through spatial and channel masks, the initial fused feature size is 128×128×256. Since the gated activation operation mechanism will attenuate the boundary contour features of some minor cracks, during the residual correction process, high-frequency and low-frequency information belonging to the same scale level before processing is obtained from the network buffer storage area. That is, the 128×128×256 RGB context source features output by the first stream and the wavelet detail source features of the same size 128×128×256 output by the second stream are concatenated with the initial fused features along the channel dimension and recombined into a 128×128×768 feature tensor. This stacked tensor is then input into the correction module.

[0069] Next, a 1×1 bottleneck convolutional layer with a channel dimension output set to 256 is used to perform feature interaction dimensionality reduction to remove information redundancy and compress the number of channels to 256. Then, a 3×3 standard spatial convolutional layer with padding set to 1 to keep the spatial resolution unchanged is used to extract spatial topological relationships. The output of this layer is a 128×128×256 dimension compensated residual.

[0070] The obtained compensation residuals are added element-wise with the preliminary fusion features. The residuals are used to repair the missing image detail boundaries, supplement the high-frequency contour features that are reduced due to the inter-feature layer transmission, and reduce the impact of background interference on the boundary features. The thin film feature reconstruction is completed and the features are sent to the decoder for pixel segmentation.

[0071] Next, the surface defects of the plastic film were evaluated through ablation experiments. The specific steps were as follows: a plastic film surface defect dataset was constructed, containing 5000 samples with a spatial resolution of 1024×1024 pixels, covering three common defects: scratches, bubbles, and holes. The dataset was divided into training, validation, and test sets in a ratio of 8:1:1. The hardware platform was equipped with two RTX 3090 graphics cards and 64GB of RAM. The deep learning framework used was PyTorch version 1.9. The AdamW optimizer was used during the network training phase, with an initial learning rate set to 0.0005, supplemented by a cosine annealing learning rate decay strategy. The batch size was set to 5, and the total number of iterations was set to 200. The average intersection-over-union ratio (IoU) and F1 score were used as the core evaluation metrics. Using a network consisting only of a single-stream ResNet50 backbone as the baseline model, and gradually adding the functional modules proposed in this embodiment, three model variants were constructed for ablation testing. Figure 3 The accompanying ablation experiment comparison diagram provided in this embodiment of the invention shows that the baseline network has an average crossover-union ratio (CUI) of 75.2% and an F1 score of 78.6%. After adding dual-tree complex wavelet transform and a dual-stream backbone network, the average CUI increases to 78.9% and the F1 score increases to 81.4%. After further superimposing a parallel receptive field module, the average CUI increases to 81.5% and the F1 score increases to 83.7%. After fully employing gating units, background suppression attention modules, and source-aware residual correction modules, the model achieves an average CUI of 84.6% and an F1 score of 86.9% on the test set. The results show that the introduction of dual-tree complex wavelet transform and dual-stream backbone network improves the average crossover ratio (CRR) by 3.7 percentage points, indicating that multi-channel high-frequency detail features can enhance the model's ability to represent defect edges and local abrupt changes. The parallel receptive field module further improves the average CRR by 2.6 percentage points, indicating that multi-scale receptive field features can reduce missed detections caused by large differences in defect size. The gating unit, background suppression attention module, and residual compensation correction mechanism further improve the average CRR by 3.1 percentage points, indicating that this combined mechanism can reduce background interference such as normal light spots and wrinkle artifacts, and compensate for the spatial detail loss generated during the inter-layer transfer of features, thereby improving the stability and boundary integrity of pixel-level segmentation of defects on the plastic film surface.

Claims

1. A method for detecting surface defects in plastic films based on multi-scale feature fusion, characterized in that, include: S1. Acquire the plastic film image and perform dual-tree complex wavelet transform to decompose the low-frequency sub-band as the global background prior, and integrate the high-frequency directional sub-band into multi-channel detail features. S2, Construct a dual-stream backbone network, wherein each stage of the dual-stream backbone network has a built-in parallel receptive field module composed of multiple convolutional branches with different dilatancy rates. S3, establish a top-down fusion path, generate temporary spatial weights and channel weights matching the number of low-level feature channels to be fused, generate an initial feature map by combining the global background prior, construct a background suppression attention map, and obtain spatial weights by combining the temporary spatial weights, thereby obtaining preliminary fused features; S4. Based on the preliminary fusion features, calculate the compensation residual, thereby restoring the original details and obtaining the feature map set. Perform pixel-level classification prediction on the feature map set and output the plastic film surface defect segmentation result.

2. The method for detecting surface defects in plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The acquisition of the plastic film image and the performance of dual-tree complex wavelet transform specifically include: inputting the single-channel brightness image corresponding to the plastic film image into a preset dual-tree separation analysis branch; performing one-dimensional low-pass and high-pass filtering and downsampling operations sequentially along the horizontal and vertical directions of the image at each decomposition level; after the decomposition loop, extracting the top-level low-pass output as the low-frequency sub-band; converting the low-frequency sub-band into a real-valued low-frequency amplitude map as the global background prior; and performing real-valued processing on the complex high-frequency output coefficients generated at each level to obtain a high-frequency amplitude feature map; upsampling the high-frequency amplitude feature map to the spatial resolution of the plastic film image and then stitching it together according to the channel dimension to obtain the multi-channel detail features.

3. The method for detecting surface defects in plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The construction of the dual-stream backbone network includes: in the first stream, retaining the initial large convolutional kernels and max-pooling downsampling operations of the pre-trained convolutional neural network to establish a main feature path for extracting global contextual features; in the second stream, adopting a convolutional layer architecture of the same depth as the first stream, with the input channel receiving the multi-channel detailed features, and performing downsampling operations synchronously in both streams, maintaining the same spatial resolution at each corresponding feature level, and obtaining structural features through parallel extraction.

4. The method for detecting surface defects of plastic films based on multi-scale feature fusion according to claim 3, characterized in that, Each stage of the dual-stream backbone network has a built-in parallel receptive field module consisting of multiple convolutional branches with different dilation rates. This includes setting up three parallel convolutional branches at the same level node for feature extraction. Each branch uses a convolutional kernel of the same size to receive the same input feature map. The multi-scale feature maps extracted by the three parallel branches are spliced ​​and fused in the channel dimension, and the output is used as the feature representation of the corresponding level.

5. The method for detecting surface defects in plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The process of generating temporary spatial weights and matching channel weights with the number of low-level feature channels to be fused includes: performing global average pooling on the input high-level features to compress the spatial dimension, obtaining global channel statistics, and then sequentially passing the channel weights through a multilayer perceptron with an output dimension equal to the number of low-level feature channels to be fused and an activation function; inputting the high-level features into a two-dimensional convolutional layer, and after processing by the activation function, generating a two-dimensional matrix as the temporary spatial weights.

6. The method for detecting surface defects of plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The construction of the background suppression attention map includes: using a bilinear interpolation algorithm to proportionally enlarge the low-frequency sub-band, which serves as the global background prior, according to the length and width dimensions of the low-level features to be fused, to obtain an enlarged feature map; inputting the enlarged feature map into a sub-network composed of convolutional layers and activation functions, outputting the initial feature map, and subtracting the initial feature map from it using a unit tensor of the same size as the initial feature map to obtain the background suppression attention map, which is used to suppress the feature response values ​​of the background noise region.

7. The method for detecting surface defects of plastic films based on multi-scale feature fusion according to claim 3, characterized in that, The process of restoring the original details and obtaining the feature map set includes: extracting the source features corresponding to the same level, which are output by the first stream before fusion, and the source features output by the second stream; concatenating the two sets of source features with the preliminary fusion features in the channel dimension to obtain composite features; inputting the composite features into a correction module for dimensionality reduction convolution and feature extraction convolution to obtain the compensation residual used to compensate for spatial and semantic information loss; and adding the compensation residual to the preliminary fusion features element-wise to complete the feature residual correction and restoration, thereby obtaining the feature map set.

8. The method for detecting surface defects of plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The process of obtaining the preliminary fusion features includes: upsampling the temporary spatial weights to the same spatial resolution as the low-level features to be fused, and multiplying them element-wise with the background suppression attention map to obtain composite spatial weights; using the composite spatial weights to perform spatial filtering on the low-level features to be fused, and using the channel weights to perform channel weighting on the low-level features to be fused to obtain the preliminary fusion features.

9. The method for detecting surface defects of plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The pixel-level classification prediction of the feature map set includes: uniformly upsampling feature maps of different resolutions in the feature map set to a preset decoding resolution, mapping them to the same number of channels, and concatenating the aligned multi-level features along the channel dimension to form a comprehensive segmentation feature map; inputting the comprehensive segmentation feature map into multiple convolutional layers and activation layers in sequence for feature decoding to generate a classification score map, selecting the category with the highest probability in the channel dimension as the category label of the corresponding pixel, and outputting the segmentation result of the plastic film surface defect.

10. The method for detecting surface defects of plastic films based on multi-scale feature fusion according to claim 1, characterized in that, The method further includes: acquiring plastic film defect images with pixel-level labeled masks as training samples, inputting the original images and the multi-channel detail features into a dual-stream backbone network respectively to obtain classification score maps of the corresponding samples; comparing the classification score maps with the pixel-level defect labels of the training samples, constructing a joint loss function composed of cross-entropy loss and Dice loss weighted together, and performing end-to-end parameter updates on the network model based on the joint loss function.