Transform-based wet tissue surface defect real-time detection method

By using a Transformer-based method that combines visible light and near-infrared image features and integrates prior information from category features, the problem of high false negative rate and poor generalization ability in wet wipe surface defect detection is solved, achieving more efficient and accurate defect detection.

CN120976205AActive Publication Date: 2025-11-18HANGZHOU GUOGUANG TOURING COMMODITY

Patent Information

Application Number
CN202511369993.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-11-18
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Traditional machine vision-based methods for detecting surface defects in wet wipes suffer from high false negative rates and poor generalization ability, especially due to differences in surface defects caused by different types of wet wipes and complex optical noise caused by moisture.

Method used

A Transformer-based approach is adopted, which combines category feature priors with multispectral feature fusion. By utilizing feature extraction and cross-attention mechanisms from visible light and near-infrared images, fused features are generated to suppress reflective noise, enhance real defect features, and perform defect detection based on wet wipe category information.

Benefits of technology

It improves the accuracy and efficiency of surface defect detection on wet wipes, reduces interference from irrelevant information, and enhances the generalization ability and robustness of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976205A_ABST
    Figure CN120976205A_ABST
Patent Text Reader

Abstract

The invention discloses a Transform-based wet tissue surface defect real-time detection method. The method comprises the following steps: acquiring category information and a detection image of a to-be-detected wet tissue; the detection image comprises a visible light image and a near-infrared image; according to the category information of the to-be-detected wet tissue, acquiring category features through a preset category feature information base, and encoding the category features to generate category feature vectors; performing feature extraction on the visible light image and the near-infrared image to obtain visible light features and near-infrared features; based on the category feature vector, dynamically fusing the visible light feature and the near-infrared feature through a cross attention mechanism to generate a fused feature; and based on the fusion features, generating a wet tissue surface defect detection result through a preset defect detection model. According to the method, the accuracy and efficiency of detecting the surface defects of different types of wet tissues can be effectively improved by combining the category feature prior with the visible light feature and near-infrared feature fusion method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and in particular to a real-time detection method for surface defects on wet wipes based on Transformer. Background Technology

[0002] As a daily hygiene product, the surface defects of wet wipes (such as stains, damage, wrinkles, etc.) directly affect product quality and user experience. Therefore, defect detection of wet wipes is an indispensable step.

[0003] Unlike other object surface inspections, such as fabrics and packaging bags, wet wipes have surface moisture that causes localized reflections, creating complex optical noise, such as light spots and brightness fluctuations caused by water films. These noises are highly similar to the visual characteristics of real defects. Furthermore, because different types of wet wipes have different surface defects, traditional machine vision-based inspection methods still suffer from high false negative rates and poor generalization ability when dealing with surface defects on wet wipes. Summary of the Invention

[0004] The purpose of this application is to provide a real-time detection method for surface defects of wet wipes based on Transformer, which improves the accuracy and efficiency of detecting surface defects of different types of wet wipes by combining prior class features with multispectral feature fusion.

[0005] Firstly, this application provides a real-time detection method for surface defects of wet wipes based on Transformer, employing the following technical solution: Acquire the category information and detection images of the wet wipes to be tested; the detection images include visible light images and near-infrared images; Based on the category information of the wet wipes to be tested, category features are obtained through a preset category feature information database, and the category features are encoded to generate a category feature vector; Feature extraction was performed on visible light images and near-infrared images respectively to obtain visible light features and near-infrared features; Based on category feature vectors, visible light features and near-infrared features are dynamically fused through a cross-attention mechanism to generate fused features; Based on fusion features, the system generates surface defect detection results for wet wipes using a pre-set defect detection model.

[0006] Through the above technical solutions, visible light features can provide the color and texture details of defects, while near-infrared features can provide anti-reflective structure and light transmission information. The cross-attention mechanism of Transformer can make the two features complementary, ultimately suppressing reflective noise and enhancing the real defect features. At the same time, using wet wipe category information as prior information, the defect detection model can focus on specific defect features according to the wet wipe category, reducing interference from irrelevant information, thereby improving the accuracy and generalization ability of wet wipe surface defect detection.

[0007] Optionally, after obtaining the category information and detection image of the wet wipe to be tested, the process includes: Based on the category information of the wet wipes to be tested, the corresponding preprocessing parameter table is obtained through a preset category parameter database; Perform general preprocessing on the detected images; Based on the preprocessing parameter table, the detected image is preprocessed specifically.

[0008] Optionally, the dynamic fusion of visible light features and near-infrared features based on the category feature vector, using a cross-attention mechanism, to generate fused features includes: Generate a category prior mapping matrix based on the category feature vectors; Based on the category prior mapping matrix, visible light features and near-infrared features are weighted to generate weighted visible light features and weighted near-infrared features; Based on weighted visible light features and weighted near-infrared features, enhanced visible light features and enhanced near-infrared features are generated through a cross-attention mechanism; Enhanced visible light features and enhanced near-infrared features are fused to generate a fused feature.

[0009] Optionally, the category prior mapping matrix includes a spatial mapping matrix and a channel mapping matrix. The step of weighting visible light features and near-infrared features based on the category prior mapping matrix to generate weighted visible light features and weighted near-infrared features includes: Based on the spatial mapping matrix and the channel mapping matrix, the visible light features and near-infrared features are weighted in the first layer to generate the first weighted visible light features and the first weighted near-infrared features. Generate a dynamic weight map based on category feature vectors; Based on the dynamic weight map, the first weighted visible light feature and the first weighted near-infrared feature are weighted in the second layer to generate weighted visible light feature and weighted near-infrared feature.

[0010] Optionally, generating a dynamic weight map based on the category feature vector includes: Based on visible light characteristics, the central and edge regions are determined by pre-setting the proportion of the edge region; Generate basic masks for the central region and the edge region respectively, and denote them as the central basic mask and the edge basic mask; The category feature vector is parsed to obtain the weights of the central region and the edge region. Based on the weights of the central region and the edge region, the central base mask and the edge base mask are weighted and summed to generate a dynamic weight map.

[0011] Optionally, the fusion of enhanced visible light features and enhanced near-infrared features to generate fused features includes: Based on the category information of the wet wipes to be tested, category preference constants are obtained through a pre-set category preference information database; Similarity calculations are performed on enhanced visible light features and enhanced near-infrared features to obtain feature similarity; The feature similarity is normalized, and feature bias coefficients are generated using the class preference constant. The weighted fusion coefficients are obtained based on the feature bias coefficients and the category preference constants. Based on the weighted fusion coefficient, the enhanced visible light features and enhanced near-infrared features are weighted and fused to generate fused features.

[0012] Optionally, the step of generating surface defect detection results for wet wipes based on fusion features and a preset defect detection model includes: Based on fusion features, the test results of surface defects of candidate wet wipes are generated by using a preset defect detection model. Based on the category information of the wet wipes to be tested, a dynamic confidence threshold is generated; Based on the surface defect detection results of candidate wet wipes, the surface defect detection results of wet wipes are generated through a dynamic confidence threshold.

[0013] Optionally, a dynamic confidence threshold is generated based on the category information of the wet wipes to be tested, including: Based on the category information of the wet wipes to be tested, the corresponding historical test data is obtained through a preset historical test database. The historical test data includes defect test data and real labeling data. Set a confidence threshold sequence and arrange them in ascending order; Traverse the confidence threshold sequence and calculate the recall rate at each confidence threshold based on defect detection data and real labeled data. After the traversal is completed, the minimum confidence threshold is obtained by using the preset target recall rate, and it is denoted as the dynamic confidence threshold.

[0014] Secondly, this application provides a Transformer-based real-time detection system for surface defects in wet wipes, comprising: The information acquisition module 101 is used to acquire the category information and detection images of the wet wipes to be tested; the detection images include visible light images and near-infrared images; The category feature vector generation module 102 is used to obtain category features based on the category information of the wet wipes to be tested through a preset category feature information library, and encode the category features to generate a category feature vector. The image feature extraction module 103 is used to extract features from visible light images and near-infrared images respectively, and obtain visible light features and near-infrared features; The image feature fusion module 104 is used to dynamically fuse visible light features and near-infrared features based on category feature vectors through a cross-attention mechanism to generate fused features; The detection result generation module 105 is used to generate the detection results of surface defects of wet wipes based on the fusion features and through the preset defect detection model.

[0015] Thirdly, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above for a Transformer-based real-time detection method for surface defects in wet wipes.

[0016] In summary, this application first uses category features as prior guidance and employs a cross-attention mechanism to fuse visible light and near-infrared features, which can effectively improve the accuracy and efficiency of surface defect detection for different types of wet wipes. Furthermore, when fusing visible light and near-infrared features, a weighted fusion coefficient is added, allowing the final fused features to further enhance the consistency and complementarity of the two modalities, thus improving the robustness of the defect detection model in complex scenarios. In addition, when generating defect detection results through the defect detection model, the detection results are filtered using a dynamic confidence threshold, based on the different defect judgment criteria for different types of wet wipes, making the final defect detection results more aligned with actual needs. Attached Figure Description

[0017] Figure 1 This is a flowchart of a real-time detection method for surface defects of wet wipes based on Transformer, provided in an embodiment of this application. Figure 2 This application provides a flowchart of a process for dynamically fusing visible light features and near-infrared features through a cross-attention mechanism to generate fused features. Figure 3 This is a flowchart of generating a dynamic weight map based on category feature vectors, provided in an embodiment of this application. Figure 4 This is a flowchart provided in an embodiment of the present application for fusing enhanced visible light features and enhanced near-infrared features to generate fused features; Figure 5 This is a schematic diagram of a real-time detection system for surface defects of wet wipes based on Transformer, provided in an embodiment of this application. Detailed Implementation

[0018] The following is in conjunction with the appendix Figure 1 -Appendix Figure 5 This application will be described in further detail below.

[0019] This application provides a real-time detection method for surface defects on wet wipes based on Transformer, see [link to relevant documentation]. Figure 1 This includes the following steps: S100: Obtain the category information and detection image of the wet wipe to be tested.

[0020] S200. Based on the category information of the wet wipes to be tested, obtain the category features through a preset category feature information database, and encode the category features to generate a category feature vector.

[0021] S300: Extract features from visible light images and near-infrared images respectively to obtain visible light features and near-infrared features.

[0022] S400: Based on the category feature vector, visible light features and near-infrared features are dynamically fused through a cross-attention mechanism to generate fused features.

[0023] S500 generates surface defect detection results for wet wipes based on fusion features and a preset defect detection model.

[0024] The detected images include visible light images and near-infrared images. Visible light images correspond to the wavelengths that the human eye can perceive (approximately 400-700nm), and can intuitively reflect information such as the color, shape, and texture of an object. They are the most basic image source for defect detection. By capturing the visible light reflected by the object (after being illuminated by natural or artificial light sources) through a camera, and focusing it onto the photosensitive element through an optical lens, a color or grayscale image can be generated. Color images are preferred here because color information needs to be preserved.

[0025] Near-infrared images correspond to the 700-1100nm wavelength band. The absorption characteristics of light in this band for moisture and organic matter are significantly different from those of visible light. This makes it suitable for weakening reflections and highlighting translucency defects such as semi-transparent impurities and fiber damage in wet wipes.

[0026] In this embodiment, the detection image of the wet wipe to be tested is first acquired. Visible light image and near-infrared image can be acquired simultaneously. For example, the wet wipe to be tested is placed on a detection platform (such as a conveyor belt). The visible light source and the near-infrared source are arranged coaxially or symmetrically to ensure that the irradiation area is consistent. The visible light camera and the near-infrared camera are controlled to take pictures simultaneously by the same trigger signal, so that the visible light image and the near-infrared image of the wet wipe to be tested can be acquired. It should be noted that the resolution of the visible light camera and the near-infrared camera is set to be the same to ensure the accuracy of subsequent image registration.

[0027] In addition, considering that different types of wipes differ in raw materials, processes, and functions, the characteristics of defects will vary. For example, baby wipes are usually softer and thinner, and may contain moisturizing ingredients, making them prone to defects such as fiber shedding, localized thinning (uneven light transmission), and residual spots. Medical disinfectant wipes may be made of non-woven fabric, which is more durable and must meet sterility requirements. Defects are often edge damage, wrinkles, and foreign object contamination (such as hair or particles). Makeup remover wipes may contain makeup remover and are more moist, making them prone to uneven liquid distribution (localized over-wet / over-dry) and dark spots on the surface caused by liquid penetration.

[0028] Therefore, in this embodiment of the application, before obtaining the detection image of the wet wipe to be tested, the category information of the wet wipe to be tested will be obtained first.

[0029] Because different types of wet wipes exhibit different surface defects, the type of wet wipes is used as supplementary information to help detect surface defects in order to improve the accuracy and efficiency of wet wipe surface defect detection.

[0030] Therefore, in this embodiment of the application, the category features are obtained by means of a preset category feature information library based on the category information of the wet wipe to be tested, and the category features are encoded to generate a category feature vector.

[0031] The preset category feature information database stores different categories of wet wipes and their corresponding category features. The category features are divided into four dimensions: material properties, liquid content characteristics, defect tendency, and interference patterns. Material properties include thickness, fiber density, and flexibility; liquid content characteristics include liquid content and liquid type; defect tendency includes high-probability defect types and high-incidence areas; and interference patterns refer to factors that interfere with defect detection. For example, baby wipes contain moisturizing ingredients and are reflective; medical wipes contain alcohol and have a special near-infrared reflectivity.

[0032] For example, the category feature information stored in the preset category feature information database is as follows: {"Baby Wipes":{"Material Properties":{"Thickness (mm)":0.08,"Fiber Density (Low / Medium / High)":"Low","Flexibility (1-5)":5},"Liquid Content":{"Liquid Content (High / Medium / Low)":"High","Liquid Type":"Moisturizing Ingredients"},"Defect Tendency":{"High Probability Defects":["Fiber Shedding","Local Over-wetting"],"High-Incidence Area (Center / Edge)":"Center"},"Interference Pattern":{"Reflective Intensity (1-5)":4,"Texture Noise (1-5)":3}},"Medical Wipes":{...},…} First, based on the category of the wet wipes to be tested, all corresponding attributes, i.e., category features, can be obtained from the preset category feature information database. Then, the category features are encoded. The purpose of encoding is to transform the category features of the wet wipes into quantitative features that the detection model can understand. Non-numerical features are converted into numerical values ​​through the set quantification rules. For example, discrete attributes are encoded using one-time heat, such as "contains moisturizing ingredients" (1=yes, 0=no); grade attributes are quantified, such as fiber density: {"low":0.2,"medium":0.5,"high":0.8}, liquid content: {"high":0.9,"medium":0.5,"low":0.2}; some values ​​are set by the user, such as the weight of the center area and the weight of the edge area: {"center":[0.8,0.2],"edge":[0.2,0.8]}, and the high probability of defects is set to 0.8; the values ​​in the quantification rules can be set according to the actual situation, and this application does not impose specific limitations.

[0033] Finally, the quantization results are concatenated into a vector, and the generated vector is denoted as the category feature vector. Taking baby wipes as an example, based on the category features of baby wipes and through the preset quantization rules, the generated category feature vector is: P=[0.08,0.2,1.0,0.9,1,0.8,0.6,0.8,0.2,0.8,0.6] (corresponding to: thickness, fiber density, flexibility, liquid content, moisturizing ingredients, fiber shedding probability, local over-wetting probability, central area weight, edge area weight, reflectivity, and texture noise).

[0034] Because raw visible light and near-infrared images may have problems such as spatial misalignment, uneven illumination, and noise interference, directly extracting features will affect the representation effect of the features. Therefore, after acquiring the detection image, it is necessary to perform preprocessing operations on the detection image, specifically including the following steps: S110. Perform general preprocessing on the detected image.

[0035] S120. Based on the category information of the wet wipes to be tested, obtain the corresponding preprocessing parameter table through a preset category parameter database.

[0036] S130. Based on the preprocessing parameter table, perform dedicated preprocessing on the detection image.

[0037] The preset category parameter database stores preprocessing parameter tables corresponding to different categories of wet wipes. The preprocessing parameter tables are exclusive preprocessing parameters set for the characteristics of each type of wet wipe. The preprocessing parameters are obtained by collecting a large number of sample images for each type of wet wipe, manually annotating defects and interference areas, statistically analyzing features such as texture direction and interference area size, testing the impact of different parameters on defect detection accuracy, and selecting the optimal value.

[0038] As mentioned above, the category characteristics of wet wipes are divided into four dimensions: material properties, liquid-containing characteristics, defect tendency, and interference mode. Therefore, the preprocessing parameters are also generated for these four feature dimensions, which are defined as texture feature parameters, spectral processing parameters, defect concern area parameters, and interference repair parameters, respectively. For different wet wipe categories, appropriate preprocessing parameters are set for each feature dimension.

[0039] In this embodiment of the application, the detection image is first subjected to general preprocessing, which is a preprocessing operation that needs to be performed on all detection images of wet wipes. General preprocessing includes image registration and band normalization.

[0040] Image registration primarily uses SIFT feature matching or subpixel-level alignment algorithms to ensure that the spatial positions of the visible light and near-infrared images correspond perfectly, thus avoiding feature misalignment caused by shooting angle deviations. Band normalization involves grayscale stretching of the near-infrared image (amplifying the difference between dark areas and the background caused by moisture absorption) and adaptive histogram equalization of the visible light image (suppressing overexposed areas caused by reflections). Then, through standardization or normalization operations, the pixel values ​​of the two images are mapped to the same range.

[0041] Because the category characteristics of wet wipes (such as material, surface morphology, liquid content, etc.) cause images to exhibit unique interference patterns, conventional image preprocessing alone cannot achieve targeted optimization of the detection images. Targeted processing based on the category characteristics of wet wipes is also required to enhance defect features and suppress category-specific noise.

[0042] Therefore, after performing general preprocessing on the detection image, based on the category information of the wet wipe to be tested, the corresponding preprocessing parameter table is obtained through a preset category parameter database to perform specific preprocessing on the detection image.

[0043] For example, regarding the texture feature parameters of material properties, baby wipes (thin non-woven fabric) have dense surface fiber textures that easily conceal minor defects such as "fiber shedding." Therefore, directional Gaussian filtering (setting the filtering direction along the fiber direction) is required to suppress lateral texture noise while preserving longitudinal fiber break edges (defect features). This means first determining the main fiber direction (such as 0° or 90°) through edge detection, and then adjusting the direction of the filtering kernel to match it, reducing the occlusion of defects by the texture. Texture feature parameters can be set accordingly, for example: filtering direction = 0, filtering size = 3×3.

[0044] After preprocessing the detection image, feature extraction can be performed on the detection image, that is, feature extraction can be performed on the visible light image and the near-infrared image separately. For example, a dual-branch Transformer encoder can be used to process the visible light image and the near-infrared image separately, and the obtained image features are recorded as visible light features and near-infrared features respectively.

[0045] Because the moisture on the surface of a wet wipe causes localized reflections and creates complex optical noise, such as light spots and brightness fluctuations caused by water films, these noises are highly similar to the visual characteristics of real defects. Therefore, multispectral images, namely visible light images and near-infrared images, can be used to achieve complementary representation of features. By utilizing the differences in response to moisture and defects in different wavelengths, the true defects on the surface of the wet wipe can be identified.

[0046] Therefore, after obtaining visible light features and near-infrared features, the cross-attention mechanism can be used to achieve the fusion of the two. Visible light features retain the color and texture details of the defects; near-infrared features can highlight the true outline of water-soluble defects and reduce reflective interference. The fused feature generated after the two are fused can suppress reflective noise and enhance the true defect features.

[0047] In addition, as mentioned above, the surface defects of different types of wet wipes exhibit different characteristics. In order to improve the accuracy and efficiency of wet wipe surface defect detection, the wet wipe category information is used as auxiliary information to help detect wet wipe surface defects, and a category feature vector is generated based on the category information of the wet wipe to be tested.

[0048] Therefore, when fusing visible light features and near-infrared features through the cross-attention mechanism, the category feature vector is used as prior information to guide the fusion weights of visible light features and near-infrared features to tilt towards the core defect features of that category of wet wipes.

[0049] Specifically, see Figure 2 Based on the category feature vector, visible light features and near-infrared features are dynamically fused through a cross-attention mechanism to generate fused features, including the following steps: S410. Generate a category prior mapping matrix based on the category feature vector.

[0050] S420. Based on the category prior mapping matrix, weighted visible light features and near-infrared features are generated to produce weighted visible light features and weighted near-infrared features.

[0051] S430 generates enhanced visible light features and enhanced near-infrared features based on weighted visible light features and weighted near-infrared features through a cross-attention mechanism.

[0052] S440: The enhanced visible light features and enhanced near-infrared features are fused to generate a fused feature.

[0053] Among them, the category prior mapping matrix is ​​used to represent the mapping weights that match the dimensions of image features. Its function is to transform the wet wipe category information into a preference guide for image features. For example, baby wet wipes prioritize the central area, while medical wet wipes prioritize the edge area.

[0054] In this embodiment, a category prior mapping matrix is ​​first generated based on the category feature vector. The category prior mapping matrix includes a spatial mapping matrix. and channel mapping matrix Space mapping matrix Used to assign weights to each spatial location in the feature map, reinforcing category-related regions; channel mapping matrix This is used to assign weights to each channel of the feature map, thereby strengthening the category-related feature channels.

[0055] Let the characteristics of visible light be denoted as Near-infrared characteristics are , , The height and width of the image feature map, The number of channels is and the category feature vector is . , dimension As in the example above, the dimension is 11.

[0056] First, the class vector is processed through fully connected layers or convolutional layers. Mapping to spatial dimensions, and then expanding by those dimensions, can generate a spatial mapping matrix. , which means assigning weights (0~1) to each spatial location (H×W pixels) of the feature map, highlighting the category-related regions, and single-channel means sharing the same spatial weights across all feature channels.

[0057] Similarly, the category feature vectors are processed through a fully connected layer. Mapping to the channel dimension generates a channel mapping matrix. This means that each feature channel is assigned a weight value (0~1), and the spatial dimension of 1×1 means that the same channel weight is shared for all spatial locations.

[0058] Once the spatial mapping matrix and channel mapping matrix are determined, the visible light features and near-infrared features can be weighted to generate weighted visible light features and weighted near-infrared features.

[0059] Specifically, based on the category prior mapping matrix, visible light features and near-infrared features are weighted to generate weighted visible light features and weighted near-infrared features, including the following steps: S421. Based on the spatial mapping matrix and the channel mapping matrix, perform a first-level weighting on the visible light features and near-infrared features respectively to generate a first-weighted visible light feature and a first-weighted near-infrared feature.

[0060] S422. Generate a dynamic weight map based on the category feature vector.

[0061] S423. Based on the dynamic weight map, perform a second layer of weighting on the first weighted visible light feature and the first weighted near-infrared feature respectively to generate weighted visible light feature and weighted near-infrared feature.

[0062] Since the defect features of different types of wet wipes differ in spatial distribution (e.g., the damage of disinfectant wipes is mostly distributed in the edge area) and channel response (e.g., the fiber density channel is more important for cotton wipes in near-infrared images), the category information can be transformed into a feature attention template through spatial mapping matrix and channel mapping matrix, so that the model can focus on the key areas and feature dimensions of the category.

[0063] First, through the spatial mapping matrix and channel mapping matrix The visible light characteristics were analyzed separately. and near-infrared features Perform the first layer of weighting, and denot the weighted visible light features and near-infrared features as the first weighted visible light features, respectively. and first weighted near-infrared features .

[0064] Since the spatial mapping matrix is ​​essentially a continuous weight map generated by convolution or fully connected layers, it may not be able to accurately lock the pixel range of the central region, meaning that class preference cannot be directly bound to the pixel range of the feature map.

[0065] Therefore, it will also focus precisely on the category characteristics of the wet wipes to be tested, that is, generate a dynamic weight map based on the category feature vector, so as to further weight the feature map. The so-called dynamic weight map is equivalent to a spatial attention mask, which can be directly applied to the feature map to further amplify the effective features of key areas and weaken the interference of irrelevant areas.

[0066] Specifically, see Figure 3Based on the category feature vectors, a dynamic weight map is generated, including the following steps: S4221. Based on visible light characteristics, determine the central region and the edge region by pre-setting the edge region proportion.

[0067] S4222. Generate basic masks for the central region and the edge region respectively, denoted as the central basic mask and the edge basic mask.

[0068] S4223. Analyze the category feature vector to obtain the weights of the central region and the edge region.

[0069] S4224. Based on the weights of the central region and the edge region, the central base mask and the edge base mask are weighted and summed to generate a dynamic weight map.

[0070] First, the size of the feature map needs to be determined, that is, the size of the visible light feature or the near-infrared feature. As mentioned above, the feature map size is H×W. Then, by presetting the edge region ratio, the edge width can be determined. Based on the edge width, the central region and the edge region can be determined.

[0071] Let the preset edge region ratio be . The edge width is ,but It can be represented as: in, It takes the integer part. This ensures that the edge width is at least 1 pixel.

[0072] For example, if the feature map size is 16×16, the preset edge region ratio typically ranges from 0.05 to 0.2. The specific value depends on the feature map size and the type of wet wipes, and can be flexibly chosen according to the actual situation. For example, if the value is 0.15, the edge width can be calculated. =2, so the size of the central region is 12×12.

[0073] After determining the central and edge regions, basic masks can be generated for the central and edge regions respectively. The basic mask is a binary mask based solely on geometric rules (1 represents the region, and 0 represents the non-region). By creating a blank mask, a central basic mask can be generated based on the central region pixel = 1 and the edge region = 0; an edge basic mask can be generated based on the edge region pixel = 1 and the central region = 0.

[0074] Then, the basic mask is fused by weighting according to category preferences. That is, the category feature vector is parsed to obtain the weights of the central region and the edge region. Finally, the central basic mask and the edge basic mask are weighted and summed according to the weights of the central region and the edge region. The resulting dynamic mask is called the dynamic weight map.

[0075] Taking the category feature vector of baby wipes mentioned above as an example, the weight of the central region (0.8) and the weight of the edge region (0.2) can be obtained by parsing. By multiplying the central base mask and the edge base mask by the corresponding region weights respectively, and then summing them, a dynamic mask covering the entire feature map can be obtained, which is the dynamic weight map.

[0076] After determining the dynamic weight map, a second layer of weighting is applied to the first weighted visible light feature and the first weighted near-infrared feature. For example, by using element-wise multiplication, the dynamic weight map is used to weight each spatial location of the first weighted visible light feature and the first weighted near-infrared feature, thereby generating the weighted visible light feature and the weighted near-infrared feature.

[0077] First, feature selection and preliminary localization are completed through spatial mapping matrix and channel mapping matrix, such as eliminating interference from low-relevance edge areas. Then, precise enhancement and region locking are completed through dynamic weight map, such as further increasing the pixel weight of the central area, further amplifying the effective features of the central area, and weakening the interference of the edge area. This allows the role of category prior to be deepened from global trend to pixel-level precise control, resulting in better detection performance for scenarios where defects are concentrated in specific areas, such as fiber shedding in the center of baby wipes or localized over-wetting scenarios.

[0078] After obtaining the weighted visible light features and weighted near-infrared features, enhanced visible light features and enhanced near-infrared features can be generated based on the weighted visible light features and weighted near-infrared features through a cross-attention mechanism.

[0079] The core of cross-attention is to make visible light features and near-infrared features pay attention to each other's important regions. For example, color anomalies in visible light may correspond to moisture anomalies in near-infrared light. Since the weighted features have been incorporated into the category prior, the generated weights are more targeted.

[0080] Specifically, firstly, the weighted visible light features and weighted near-infrared features Projecting these values ​​as query (Q), key (K), and value (V), we use a weighted approach for the query (Q) of visible light features and a weighted approach for the key (K) and value (V) of near-infrared features. We calculate the attention score between Q and K, then normalize the attention score, for example, using the softmax function. Finally, we sum the normalized attention score with V in a weighted manner to obtain the cross-attention of visible light features on near-infrared features (i.e., the information in near-infrared features related to visible light features), denoted as […]. .

[0081] Then according to By enhancing the weighted visible light features, we can obtain enhanced visible light features. .

[0082] Similarly, by using the query (Q) of weighted near-infrared features and the key (K) and value (V) of weighted visible light features, the cross-attention of near-infrared features to visible light features can be obtained, denoted as Then according to By enhancing the weighted near-infrared features, enhanced near-infrared features can be obtained. .

[0083] Finally, the enhanced visible light features and enhanced near-infrared features are fused to obtain the final fused features.

[0084] Considering the different contributions of visible light features and near-infrared features to the detection of surface defects in wet wipes, a weighted fusion coefficient is set when fusing enhanced visible light features and enhanced near-infrared features. This allows the final fused features to more flexibly balance the feature information of the two modes, which helps to improve the robustness of surface defect detection in wet wipes.

[0085] Specifically, see Figure 4 The enhanced visible light features and enhanced near-infrared features are fused to generate a fused feature, including the following steps: S4401. Based on the category information of the wet wipes to be tested, obtain the category preference constant through a preset category preference information database.

[0086] S4402. Calculate the similarity between the enhanced visible light features and the enhanced near-infrared features to obtain the feature similarity.

[0087] S4403. Normalize the feature similarity and generate feature bias coefficients using the category preference constant.

[0088] S4404. Obtain the weighted fusion coefficient based on the feature bias coefficient and the category preference constant.

[0089] S4405. Based on the weighted fusion coefficient, the enhanced visible light features and enhanced near-infrared features are weighted and fused to generate fused features.

[0090] The preset category preference information database stores category preference constants for different categories of wet wipes. The category preference constant α is used to characterize the dependence of different categories of wet wipes on visible light and near-infrared features. The closer α is to 1, the higher the dependence on visible light characteristics; the closer α is to 0, the higher the dependence on near-infrared characteristics. For example, medical wipes have a higher probability of edge breakage and are more dependent on visible light characteristics, so α is set higher, for example, 0.7. Baby wipes, due to the reflection interference caused by moisture, are more dependent on near-infrared characteristics, so α is set lower, for example, 0.3.

[0091] The feature bias coefficient β is used to bias the contribution of features based on the degree of correlation between visible light features and near-infrared features. It is measured by the similarity between features. When the similarity between two features is high, it indicates that the information complementarity is weak and the consistency is strong, and the contributions of the two features are in a balanced state. When the similarity is low, it indicates that the information complementarity is strong, and the feature bias coefficient guides the feature to tilt towards the category-biased feature.

[0092] Feature bias coefficient , can be represented as: Where s is the similarity score between enhanced visible light features and enhanced near-infrared features. The similarity between enhanced visible light features and enhanced near-infrared features is calculated and normalized to [0, 1], which is the similarity score s; 1-s is used to measure the degree of dependence on category preference.

[0093] When the similarity is extremely low (s=0): β=α, completely dependent on category preference; when the similarity is extremely high (s=1): β=0.5, the contributions of the two features are balanced.

[0094] By combining the category preference constant α and the feature bias coefficient β, a weighted fusion coefficient can be generated. , It can be represented as: in, This setting controls the contribution ratio of category preference and feature bias. It can be set according to the actual situation. For example, setting it to 0.5 means that the contribution ratios of category preference and feature bias are consistent. The closer a value is to 1, the more important the category preference becomes. The closer a value is to 0, the more important the feature bias becomes.

[0095] Finally, based on the weighted fusion coefficients, the enhanced visible light features and enhanced near-infrared features can be weighted and fused to generate fused features. , It can be represented as: The final fused feature contains key information from visible light features verified by near-infrared features and key information from near-infrared features verified by visible light features. It also retains the original feature information of both modalities, which can reflect the correlation between the two modal features and retain their uniqueness. Furthermore, by dynamically allocating weighted fusion coefficients, the final fused feature can further improve the consistency and complementarity of the two modal features, thereby helping to improve the robustness of the defect detection model in the face of complex scenarios.

[0096] After obtaining the fusion features, the surface defect detection results of the wet wipes can be generated through the preset defect detection model.

[0097] Specifically, based on fusion features, and through a pre-set defect detection model, the surface defect detection results of the wet wipes are generated, including the following steps: S510. Based on the fusion features, generate the surface defect detection results of the candidate wet wipes through the preset defect detection model.

[0098] S520. Generate a dynamic confidence threshold based on the category information of the wet wipes to be tested.

[0099] S530. Based on the candidate wet wipe surface defect detection results, generate wet wipe surface defect detection results through dynamic confidence threshold.

[0100] The preset defect detection model is essentially a lightweight target detection model. For example, it uses a lightweight detection architecture based on Transformer, with various wet wipe surface defects as the target categories. It collects sample images of various types of wet wipes (including visible light images and near-infrared images), forms a dataset through defect annotation, and trains a wet wipe surface defect detection model. The model input is fused features, and the output is the defect detection result, which includes defect category, bounding box coordinates, and confidence score.

[0101] For example, the defect detection model adopts a "Transformer decoder + detection head" structure, which forms an end-to-end process with the front-end feature fusion module. The specific structure is as follows: The Transformer decoder receives fused features as memory features and inputs learnable defect query vectors. Through self-attention and cross-attention, the query vectors interact with the fused features to locate the spatial position and feature information of potential defects. The defect query vectors can be understood as defect detectors. Before model training, the defect query vectors are randomly initialized. During training, the model learns defect samples, allowing each defect query vector to learn typical features of a certain type of defect (such as color, texture, etc.). During actual detection, the defect query vectors actively search for matching defect regions in the fused features.

[0102] The detection head consists of multiple fully connected or convolutional layers, which map the defect query vector output by the Transformer decoder to specific detection results, including: defect category (such as "stain", "damage", "uneven fiber", etc.), bounding box coordinates (such as (x1,y1,x2,y2)), and confidence (the reliability score of the prediction result). The detection result at this time can be recorded as the candidate wet wipe surface defect detection result.

[0103] The next step is to filter out low-confidence results through threshold screening to form the final wet wipe surface defect detection results. For example, if the threshold is set to 0.4, only candidate wet wipe surface defect detection results with a confidence level ≥ 0.4 will be retained.

[0104] Because the defect judgment criteria for different types of wipes are different (for example, "slight lint" in baby wipes may be considered a defect, while the same type of lint in regular wipes can be ignored), and the probability of defects occurring in different types of wipes is also different (for example, baby wipes are prone to "fiber shedding" and "uneven humidity", while makeup remover wipes are more prone to "micro-holes").

[0105] Therefore, a suitable confidence threshold is set for different wet wipes, which is called the dynamic confidence threshold. That is, based on the category information of the wet wipe to be tested, the dynamic confidence threshold is first obtained, and then the surface defect detection results of the candidate wet wipes are filtered through the dynamic confidence threshold to obtain the final surface defect detection results of the wet wipes.

[0106] Specifically, based on the category information of the wet wipes to be tested, a dynamic confidence threshold is generated, including the following steps: S521. Based on the category information of the wet wipes to be tested, obtain the corresponding historical test data through a preset historical test database.

[0107] S522. Set the confidence threshold sequence and arrange them in ascending order.

[0108] S523. Traverse the confidence threshold sequence and calculate the recall rate at each confidence threshold based on the defect detection data and the real labeled data.

[0109] S524. After the traversal is completed, the minimum confidence threshold is obtained by using the preset target recall rate, and it is denoted as the dynamic confidence threshold.

[0110] The preset historical testing database stores historical testing data for each category of wet wipes. The historical testing data includes defect testing data and real labeling data. The defect testing data is the surface defect testing results of the wet wipes mentioned above, and each test result includes a corresponding confidence level.

[0111] First, based on the category information of the wet wipes to be tested, the corresponding historical test data can be obtained through a preset historical test database.

[0112] Then, a confidence threshold sequence is set and arranged in ascending order. For example, the confidence threshold sequence can be set as [0.1, 0.2, 0.3, ..., 0.9]. This can be set according to the actual situation, and this application does not impose any specific limitations.

[0113] By traversing the confidence threshold sequence, the recall rate at that confidence threshold can be calculated based on the defect detection data and the real labeled data. The recall rate is the proportion of the number of correctly detected real defects to the total number of all real defects. Based on the real labeled data, the proportion of the total number of all real defects can be determined. Based on the confidence threshold, the number of correctly detected real defects can be calculated. In this way, the corresponding recall rate can be calculated according to the confidence thresholds in the confidence threshold sequence.

[0114] By setting a target recall rate, the minimum confidence threshold required to meet the target recall rate is denoted as the dynamic confidence threshold. For example, if the target recall rate is set to 95%, the recall rate is 98% when the confidence threshold is 0.1, 97% when the confidence threshold is 0.2, and 95% when the confidence threshold is 0.3. At this point, the target recall rate has been achieved, corresponding to a confidence threshold of 0.3. Therefore, the dynamic confidence threshold is 0.3.

[0115] By setting dynamic confidence thresholds, defect judgment criteria can be customized for each type of wet wipes, replacing fixed thresholds and making the final defect detection results more in line with actual needs.

[0116] This application also provides a Transformer-based real-time detection system for surface defects in wet wipes, see [link to relevant documentation]. Figure 5 The system includes: an information acquisition module 101, a category feature vector generation module 102, an image feature extraction module 103, an image feature fusion module 104, and a detection result generation module 105.

[0117] The information acquisition module 101 is used to acquire the category information and detection images of the wet wipes to be tested, including visible light images and near-infrared images.

[0118] The category feature vector generation module 102 is used to obtain category features based on the category information of the wet wipes to be tested through a preset category feature information library, and encode the category features to generate a category feature vector.

[0119] The image feature extraction module 103 is used to extract features from visible light images and near-infrared images respectively, and obtain visible light features and near-infrared features.

[0120] The image feature fusion module 104 is used to dynamically fuse visible light features and near-infrared features based on category feature vectors through a cross-attention mechanism to generate fused features.

[0121] The detection result generation module 105 is used to generate the detection results of surface defects of wet wipes based on the fusion features and through the preset defect detection model.

[0122] In this embodiment of the application, the information acquisition module 101 is specifically used to acquire the category information and detection image of the wet wipe to be tested, wherein the detection image includes a visible light image and a near-infrared image.

[0123] The category feature vector generation module 102 is specifically used to obtain category features based on the category information of the wet wipe to be tested obtained by the information acquisition module 101, through a preset category feature information database, and to encode the category features to generate a category feature vector.

[0124] The image feature extraction module 103 is specifically used to extract features from the visible light image and the near-infrared image acquired by the information acquisition module 101, respectively, to obtain visible light features and near-infrared features.

[0125] The image feature fusion module 104 is specifically used to dynamically fuse visible light features and near-infrared features based on the category feature vector generated by the category feature vector generation module 102 through a cross-attention mechanism to generate fused features.

[0126] The detection result generation module 105 is specifically used to generate the surface defect detection result of the wet wipe based on the fusion features generated by the image feature fusion module 104 and a preset defect detection model.

[0127] This application also provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed any of the above-described Transformer-based real-time detection methods for surface defects in wet wipes.

[0128] The embodiments described in this application are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the principles of this application should be included within the scope of protection of this application.

Claims

1. A method for real-time detection of surface defects in wet wipes based on Transformer, characterized in that, include: Acquire the category information and detection images of the wet wipes to be tested; the detection images include visible light images and near-infrared images; Based on the category information of the wet wipes to be tested, category features are obtained through a preset category feature information database, and the category features are encoded to generate a category feature vector; Feature extraction was performed on visible light images and near-infrared images respectively to obtain visible light features and near-infrared features; Based on category feature vectors, visible light features and near-infrared features are dynamically fused through a cross-attention mechanism to generate fused features; Based on fusion features, the system generates surface defect detection results for wet wipes using a pre-set defect detection model.

2. The method for real-time detection of surface defects in wet wipes based on Transformer according to claim 1, characterized in that, After obtaining the category information and detection image of the wet wipe to be tested, the process includes: Based on the category information of the wet wipes to be tested, the corresponding preprocessing parameter table is obtained through a preset category parameter database; Perform general preprocessing on the detected images; Based on the preprocessing parameter table, the detected image is preprocessed specifically.

3. The method for real-time detection of surface defects on wet wipes based on Transformer according to claim 1, characterized in that, The method of dynamically fusing visible light features and near-infrared features based on category feature vectors through a cross-attention mechanism to generate fused features includes: Generate a category prior mapping matrix based on the category feature vectors; Based on the category prior mapping matrix, visible light features and near-infrared features are weighted to generate weighted visible light features and weighted near-infrared features; Based on weighted visible light features and weighted near-infrared features, enhanced visible light features and enhanced near-infrared features are generated through a cross-attention mechanism; Enhanced visible light features and enhanced near-infrared features are fused to generate a fused feature.

4. The method for real-time detection of surface defects on wet wipes based on Transformer according to claim 3, characterized in that, The category prior mapping matrix includes a spatial mapping matrix and a channel mapping matrix. The step of weighting visible light features and near-infrared features based on the category prior mapping matrix to generate weighted visible light features and weighted near-infrared features includes: Based on the spatial mapping matrix and the channel mapping matrix, the visible light features and near-infrared features are weighted in the first layer to generate the first weighted visible light features and the first weighted near-infrared features. Generate a dynamic weight map based on category feature vectors; Based on the dynamic weight map, the first weighted visible light feature and the first weighted near-infrared feature are weighted in the second layer to generate weighted visible light feature and weighted near-infrared feature.

5. The method for real-time detection of surface defects on wet wipes based on Transformer according to claim 4, characterized in that, The generation of a dynamic weight map based on category feature vectors includes: Based on visible light characteristics, the central and edge regions are determined by pre-setting the proportion of the edge region; Generate basic masks for the central region and the edge region respectively, and denote them as the central basic mask and the edge basic mask; The category feature vector is parsed to obtain the weights of the central region and the edge region. Based on the weights of the central region and the edge region, the central base mask and the edge base mask are weighted and summed to generate a dynamic weight map.

6. The method for real-time detection of surface defects on wet wipes based on Transformer according to claim 3, characterized in that, The process of fusing enhanced visible light features and enhanced near-infrared features to generate a fused feature includes: Based on the category information of the wet wipes to be tested, category preference constants are obtained through a pre-set category preference information database; Similarity calculations are performed on enhanced visible light features and enhanced near-infrared features to obtain feature similarity; The feature similarity is normalized, and feature bias coefficients are generated using the class preference constant. The weighted fusion coefficients are obtained based on the feature bias coefficients and the category preference constants. Based on the weighted fusion coefficient, the enhanced visible light features and enhanced near-infrared features are weighted and fused to generate fused features.

7. The method for real-time detection of surface defects on wet wipes based on Transformer according to claim 1, characterized in that, The method of generating surface defect detection results for wet wipes based on fusion features and a preset defect detection model includes: Based on fusion features, the test results of surface defects of candidate wet wipes are generated by using a preset defect detection model. Based on the category information of the wet wipes to be tested, a dynamic confidence threshold is generated; Based on the surface defect detection results of candidate wet wipes, the surface defect detection results of wet wipes are generated through a dynamic confidence threshold.

8. The method for real-time detection of surface defects on wet wipes based on Transformer according to claim 7, characterized in that, Based on the category information of the wet wipes to be tested, a dynamic confidence threshold is generated, including: Based on the category information of the wet wipes to be tested, the corresponding historical test data is obtained through a preset historical test database. The historical test data includes defect test data and real labeling data. Set a confidence threshold sequence and arrange them in ascending order; Traverse the confidence threshold sequence and calculate the recall rate at each confidence threshold based on defect detection data and real labeled data. After the traversal is completed, the minimum confidence threshold is obtained by using the preset target recall rate, and it is denoted as the dynamic confidence threshold.

9. A real-time detection system for surface defects of wet wipes based on Transformer, characterized in that, include: The information acquisition module (101) is used to acquire the category information and detection image of the wet wipe to be tested; the detection image includes a visible light image and a near-infrared image; The category feature vector generation module (102) is used to obtain category features based on the category information of the wet wipes to be tested through a preset category feature information library, and encode the category features to generate a category feature vector; The image feature extraction module (103) is used to extract features from the visible light image and the near-infrared image respectively, and obtain visible light features and near-infrared features; The image feature fusion module (104) is used to dynamically fuse visible light features and near-infrared features based on the category feature vector through a cross-attention mechanism to generate fused features; The detection result generation module (105) is used to generate the detection results of surface defects of wet wipes based on the fusion features and through the preset defect detection model.

10. A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing a Transformer-based real-time detection method for surface defects on wet wipes as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Surface defect detection method and system based on Swin Transform

    CN116703885A

  • Metal surface corrosion detection method, system and equipment

    CN117576032A

  • Non-woven fabric breathable film intelligent manufacturing production line control system and method

    CN118778582A

  • Workpiece surface defect detection method based on visual identification

    CN118898620A

  • Wet tissue quality control system based on artificial intelligence

    CN118938847A

Cited By

  • Chip appearance defect detection system

    CN121917566A