Intelligent Identification Method for Defects on Printed Fabrics Based on Multimodal Data

Through intelligent identification methods based on multimodal data, combined with high-resolution RGB cameras, near-infrared camera equipment and hyperspectral imaging, features are extracted and fused, and defect identification is used using multimodal detection models, solving the problem of time-consuming, labor-intensive and susceptible to human factors in traditional manual detection methods, and achieving high-precision and high-efficiency defect detection on printed fabric surfaces.

CN119672484BActive Publication Date: 2025-05-27FUJIAN YUBANG TEXTILE TECH CO LTD

Patent Information

Application Number
CN202510193917.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Traditional manual testing methods are time-consuming and labor-intensive in textile production, and are susceptible to human factors, resulting in inconsistent testing results, making it difficult to meet the needs of modern manufacturing for efficient and accurate testing.

Method used

The intelligent identification method of printed fabric surface defects based on multimodal data is adopted, and images are acquired and pre-processed through high-resolution RGB cameras. Combined with near-infrared imaging and hyperspectral imaging, features are extracted and fused, and a multimodal detection model is used for fine-grained defect recognition and positioning, and an attention mechanism is added to improve the ability to capture complex features.

Benefits of technology

It realizes high-precision and high-efficiency defect detection on printed fabric surfaces, significantly improving the reliability and accuracy of the inspection, and can meet the real-time inspection requirements on high-speed production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672484B_ABST
    Figure CN119672484B_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent recognition method for defects on printed fabric surfaces based on multi-modal data. S1: Use a high-resolution RGB camera to obtain a global image of the printed fabric surface and perform preprocessing. S2: Based on an image recognition model, identify potential defect candidate regions and combine non-maximum suppression to reduce overlapping defect candidate regions. S3: Obtain an image of the defect candidate region through a near-infrared imaging device and obtain a hyperspectral image through hyperspectral imaging. S4: Extract features from the near-infrared image and hyperspectral image of the defect candidate region and perform feature fusion. S5: Based on the fused features and a multi-modal detection model, perform fine-grained defect recognition and localization on the defect candidate region. S6: Through post-processing steps, remove false positives and merge adjacent regions to improve the defect detection accuracy, and finally generate a detection report. The present invention can effectively identify and evaluate defects on printed fabric surfaces and achieve high-precision and high-efficiency detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of defect detection, and in particular to a method for intelligently identifying defects on a printed cloth surface based on multimodal data. Background Art

[0002] In modern manufacturing, especially in the field of textile production, product quality control has become one of the key factors in corporate competitiveness. As the global market's expectations for textile quality continue to increase, traditional manual inspection methods have shown their limitations. They are not only time-consuming and labor-intensive, but also easily affected by human factors, resulting in inconsistent inspection results. For this reason, intelligent and automated defect detection technology has gradually become a hot topic in research and application.

[0003] In recent years, with the rapid development of computer vision and machine learning technologies, intelligent defect recognition systems based on visual inspection have received increasing attention due to their advantages such as high efficiency, accuracy and good repeatability. Summary of the invention

[0004] In order to solve the above problems, the purpose of the present invention is to provide an intelligent identification method for printed cloth surface defects based on multimodal data, which can effectively identify and evaluate printed cloth surface defects and achieve high-precision and high-efficiency detection.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for intelligently identifying printed fabric surface defects based on multimodal data comprises the following steps:

[0007] S1: Use a high-resolution RGB camera to obtain a global image of the printed cloth surface and perform preprocessing;

[0008] S2: Based on the image recognition model, potential defect candidate areas are identified, and non-maximum suppression is combined to reduce overlapping defect candidate areas;

[0009] S3: Acquire an image of the defect candidate area through a near-infrared camera device, and acquire a hyperspectral image through hyperspectral imaging;

[0010] S4: extracting features from the near infrared image and hyperspectral image of the defect candidate area and performing feature fusion;

[0011] S5: Based on the fused features and the multimodal detection model, fine-grained defect recognition and positioning are performed on the defect candidate areas, and an attention mechanism is added to improve the ability to capture complex features of defects;

[0012] S6: Through post-processing steps, false positives are removed and adjacent areas are merged to improve the accuracy of defect detection. Finally, a test report is generated, marking the location, type and possible area of ​​the defect, so that the operator can make adjustments to the subsequent process.

[0013] Furthermore, the preprocessing is as follows:

[0014] Through the automatic white balance algorithm, the gains of the red, green and blue channels are adjusted in real time to ensure the authenticity and consistency of the colors;

[0015] Use Gaussian filter to remove image noise, enhance image signal-to-noise ratio, and use Fourier transform to remove periodic noise;

[0016] The image contrast is enhanced by histogram equalization, making defects easier to identify, and the gamma value is adjusted to optimize the overall brightness of the image, so as to better show the texture and color details of the fabric in the unsaturated color spectrum.

[0017] Furthermore, a Gaussian filter is used to remove image noise and enhance the image signal-to-noise ratio, and Fourier transform is used to remove periodic noise, as follows:

[0018] The Gaussian filter convolves the image with a Gaussian kernel:

[0019] ;

[0020] Among them, 2g+1 is the size of the Gaussian kernel, I(x,y) represents the pixel value of the input image, I′(x,y) represents the pixel value of the output image; (x,y) is the position of the pixel in the image; (i,j) is the relative offset coordinate of the Gaussian kernel; represents Gaussian kernel;

[0021] Fourier transform converts the image processed by Gaussian filtering from the spatial domain to the frequency domain, identifies and removes noise of specific frequencies, and performs a two-dimensional Fourier transform on the input image I′(x, y) to obtain the frequency domain representation F(u, v);

[0022] ;

[0023] Where M and N represent the width and height of the input image respectively; u and v are the indices of the two-dimensional frequency; is an imaginary unit;

[0024] Identify and suppress high frequencies or specific frequency bands that represent periodic noise, and set the corresponding noise frequencies to zero in the Fourier plane:

[0025] F′(u,v)=F(u,v)×H(u,v)

[0026] Where H(u,v) is the frequency domain filter;

[0027] Perform a two-dimensional inverse Fourier transform on the processed frequency domain image F′(u,v) and restore it to the spatial domain:

[0028] .

[0029] Furthermore, the image recognition model is built based on the convolutional feature extraction network, the region proposal network and the Faster R-CNN, as follows:

[0030] The convolutional feature extraction network uses the ResNet network to extract rich feature representations on the image.

[0031] The region proposal network RPN generates a set of candidate regions on the feature map through a sliding window. When each window slides to a feature point, multiple candidate region anchors of different scales and aspect ratios are generated;

[0032] RPN uses a binary classification model to determine whether each anchor contains a target and a specific defect category;

[0033] The precise boundary of the candidate region is further adjusted through regression, and the loss function is:

[0034] ;

[0035] in, Candidate area The probability value of being predicted as a certain category; Candidate area The true category label of is the predicted bounding box regression parameter; is the true bounding box regression parameter; is the normalization factor of the classification loss, is the normalization factor of the regression loss, is the loss weight coefficient, is the classification loss function; is the regression loss function:

[0036] Fast R-CNN extracts RoI features, crops fixed-size features from the convolutional feature map through RoI Pooling based on the candidate regions generated by RPN, and uses a fully connected layer to perform category discrimination and bounding box refinement on each candidate feature to obtain defect candidate regions.

[0037] Furthermore, non-maximum suppression is combined to reduce overlapping defect candidate areas, as follows:

[0038] Arrange the bounding boxes of all detection results in descending order according to their category confidence.

[0039] For a bounding box set {B k}, sorted in descending order of confidence:

[0040] L(B)={B 1 ,B 2 ,...,B k ,...,B n ∣ score(B k )≥score(B k+1 )};

[0041] From the sorted candidate box set L(B), select the box B with the highest confidence in each cycle max ;

[0042] Calculate overlap IOU:

[0043]

[0044] For all other boxes B k , if IOU(B k ,B max )>θ, θ is the preset threshold, then adjust the integral score :

[0045]

[0046] in, To control the parameters;

[0047] Boxes with scores below a preset threshold are removed and the update calculation is repeated.

[0048] Furthermore, S4 is specifically:

[0049] Standardize the input near-infrared image and hyperspectral image data, including distribution normalization for each band;

[0050] The hyperspectral image data is subjected to dimensionality reduction through principal component analysis;

[0051] CNN is applied to near-infrared images and hyperspectral image data after dimensionality reduction to extract basic feature representations;

[0052] f NIR =CNN NIR (x NIR );

[0053] f HSI =CNN HSI (x HSI );

[0054] Among them, x NIR and x HSI They are near-infrared images and hyperspectral image data after dimensionality reduction; CNN NIR and CNN HSI are the convolutional neural networks pre-trained for the corresponding images; f NIR and f HSI are the extracted near-infrared features and hyperspectral features respectively; W NIR and W HSI is the weight;

[0055] Using the joint feature stacking strategy, f NIR and f HSI Generate the fusion feature vector f by weighting u :

[0056] ;

[0057] Adding position encoding PE to the fusion feature allows the model to take advantage of the sequence order:

[0058] ;

[0059] Among them, pos refers to the position, is the dimension index, d model is the dimension of the feature.

[0060] Furthermore, the multimodal detection model is as follows:

[0061] A multimodal detection model is built based on the improved R-CNN architecture, and a custom module is used to introduce channel attention mechanism and spatial attention mechanism to highlight the important features of defects;

[0062] Channel attention emphasizes the important channels in the feature map and generates channel attention M through global average pooling and global maximum pooling c (X):

[0063] M c (X) = σ s (W 1 (δ(W 0 (F a (X)))))+W 1 (δ(W 0 (F m (X))));

[0064] Among them, σ s is the Sigmoid activation function; δ is the nonlinear activation function; W 0 and W 1 is the weight matrix;

[0065] F a (X), F m (X) are the global average pooling and global maximum pooling operations of the input feature map X respectively;

[0066] Spatial AttentionM s (X), used to emphasize the area with significant features in the feature map:

[0067] M s (X) = σ s (f 7×7 ([F a (X);F m (X)]));

[0068] Among them, f 7×7 Indicates that a 7×7 convolution kernel is used for convolution operation;

[0069] Simultaneously learn the classification task and bounding box regression task;

[0070] Combined with the output of the attention module, it generates category probabilities for defect recognition and bounding box predictions for precise positioning;

[0071] The loss function L includes the classification loss using the cross entropy loss L cls , regression loss uses smooth L1 loss L reg :

[0072] ;

[0073] in: is the classification loss, which is used to evaluate the predicted result p and the true label The difference between is the regression loss used to evaluate the predicted bounding box parameters With the real parameters The difference between is a hyperparameter; a represents the ath sample, Represents a binary switch variable that determines whether the a-th candidate box needs to calculate the regression loss of the bounding box.

[0074] Furthermore, the R-CNN architecture is improved as follows:

[0075] The region proposal network RPN predicts the objectness score and bounding box regression of the anchor box through a sliding window;

[0076] RoI pooling layer converts feature maps of candidate regions of different sizes into fixed sizes; RoI feature=pool(F,RoIs)

[0077] Where F is the input feature map, RoIs are candidate boxes generated by RPN, and RoI pooling adjusts the mapping according to the size of each RoI to make the output RoI feature size fixed;

[0078] A deep feature extraction network is used to perform a classification and regression joint training strategy to improve the detection accuracy, and fine classification and precise regression of the bounding box are performed for each RoI.

[0079] Furthermore, S6 is specifically:

[0080] Assign a confidence score to each detection result, set a confidence threshold, and filter out detection results below the threshold; use a secondary classifier to reclassify the preliminary detection results to distinguish between real targets and falsely detected targets; combine scene knowledge or location restrictions to filter out unreasonable detection results and obtain the final detection results;

[0081] According to the final test results, draw a box on the image to mark the defect location, mark the test type result and attach the test confidence level;

[0082] Perform pixel counting on the merged area to estimate the number of pixels occupied by defects, and create an automated module to generate a test report in a standard format, including the number, location, type, area and corresponding recommendations of defects.

[0083] The present invention has the following beneficial effects:

[0084] 1. The present invention can effectively identify and evaluate defects on printed fabrics, achieve high-precision and high-efficiency detection, significantly improve the reliability and accuracy of detection, and meet the real-time detection requirements on high-speed production lines;

[0085] 2. The present invention constructs an image recognition model based on a convolutional feature extraction network, a region proposal network and Faster R-CNN to efficiently and accurately detect and identify candidate defect regions in printed cloth images;

[0086] 3. The present invention can effectively fuse and extract features between different modalities through cross-modal feature fusion, and based on the improved RCNN model, it can significantly reduce the amount of redundant candidate frames to be processed by combining the RPN network, and enhance feature extraction and expression by introducing RoI pooling and attention modules, so as to achieve higher accuracy in identifying and classifying defects caused by different materials or dyes of printed cloth. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0088] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0089] refer to Figure 1 In this embodiment, a method for intelligently identifying printed fabric surface defects based on multimodal data is provided, comprising the following steps:

[0090] S1: Use a high-resolution RGB camera to obtain a global image of the printed cloth surface and perform preprocessing;

[0091] S2: Based on the image recognition model, potential defect candidate areas are identified, and non-maximum suppression is combined to reduce overlapping defect candidate areas;

[0092] S3: Acquire an image of the defect candidate area through a near-infrared camera device, and acquire a hyperspectral image through hyperspectral imaging;

[0093] S4: extracting features from the near infrared image and hyperspectral image of the defect candidate area and performing feature fusion;

[0094] S5: Based on the fused features and the multimodal detection model, fine-grained defect recognition and positioning are performed on the defect candidate areas, and an attention mechanism is added to improve the ability to capture complex features of defects;

[0095] S6: Through post-processing steps, false positives are removed and adjacent areas are merged to improve the accuracy of defect detection. Finally, a test report is generated, marking the location, type and possible area of ​​the defect, so that the operator can make adjustments to the subsequent process.

[0096] In this embodiment, the preprocessing is as follows:

[0097] The automatic white balance (AWB) algorithm adjusts the gain of the red, green and blue channels in real time to ensure color authenticity and consistency;

[0098] Use Gaussian filter to remove image noise, enhance image signal-to-noise ratio, and use Fourier transform to remove periodic noise;

[0099] The image contrast is enhanced by histogram equalization, making defects easier to identify, and the gamma value is adjusted to optimize the overall brightness of the image, so as to better show the texture and color details of the fabric in the unsaturated color spectrum.

[0100] In this embodiment, a Gaussian filter is used to remove image noise and enhance the image signal-to-noise ratio, and a Fourier transform is used to remove periodic noise, as follows:

[0101] The Gaussian filter convolves the image with a Gaussian kernel:

[0102] ;

[0103] Among them, 2g+1 is the size of the Gaussian kernel, I(x,y) represents the pixel value of the input image, I′(x,y) represents the pixel value of the output image; (x,y) is the position of the pixel in the image; (i,j) is the relative offset coordinate of the Gaussian kernel; represents Gaussian kernel;

[0104] Fourier transform converts the image processed by Gaussian filtering from the spatial domain to the frequency domain, identifies and removes noise of specific frequencies, and performs a two-dimensional Fourier transform on the input image I′(x,y) to obtain the frequency domain representation F(u,v);

[0105] ;

[0106] Where M and N represent the width and height of the input image respectively; u and v are the indices of the two-dimensional frequency; is an imaginary unit;

[0107] Identify and suppress high frequencies or specific frequency bands that represent periodic noise, and set the corresponding noise frequencies to zero in the Fourier plane:

[0108] F′(u,v)=F(u,v)×H(u,v)

[0109] Where H(u,v) is the frequency domain filter;

[0110] Perform a two-dimensional inverse Fourier transform on the processed frequency domain image F′(u,v) and restore it to the spatial domain:

[0111] .

[0112] In this embodiment, the image recognition model is constructed based on a convolutional feature extraction network, a region proposal network, and a Faster R-CNN, as follows:

[0113] The convolutional feature extraction network uses the ResNet network to extract rich feature representations on the image.

[0114] The region proposal network RPN generates a set of candidate regions on the feature map through a sliding window. When each window slides to a feature point, multiple candidate region anchors of different scales and aspect ratios are generated;

[0115] RPN uses a binary classification model to determine whether each anchor contains an object (target) and a specific defect category;

[0116] The precise boundary of the candidate region is further adjusted through regression, and the loss function is:

[0117] ;

[0118] in, Candidate area The probability value of being predicted as a certain category; Candidate area The true category label of is the predicted bounding box regression parameter; is the true bounding box regression parameter; is the normalization factor of the classification loss, is the normalization factor of the regression loss, is the loss weight coefficient, is the classification loss function; is the regression loss function:

[0119] Fast R-CNN extracts RoI features, crops fixed-size features from the convolutional feature map through RoI Pooling based on the candidate regions generated by RPN, and uses a fully connected layer to perform category discrimination and bounding box refinement on each candidate feature to obtain defect candidate regions.

[0120] In this embodiment, non-maximum suppression is combined to reduce overlapping defect candidate regions, as follows:

[0121] Arrange the bounding boxes of all detection results in descending order according to their category confidence.

[0122] For a bounding box set {B k}, sorted in descending order of confidence:

[0123] L(B)={B 1 ,B 2 ,...,B k ,...,B n ∣ score(B k )≥score(B k+1 )};

[0124] From the sorted candidate box set L(B), select the box B with the highest confidence in each cycle max ;

[0125] Calculate overlap IOU:

[0126]

[0127] For all other boxes B k , if IOU(B k ,B max )>θ, θ is the preset threshold, then adjust the integral score :

[0128]

[0129] in, To control the parameters;

[0130] Boxes with scores below a preset threshold are removed and the update calculation is repeated.

[0131] In this embodiment, S4 is specifically:

[0132] Standardize the input near-infrared image and hyperspectral image data, including distribution normalization for each band;

[0133] The hyperspectral image data is subjected to dimensionality reduction through principal component analysis;

[0134] CNN is applied to near-infrared images and hyperspectral image data after dimensionality reduction to extract basic feature representations;

[0135] f NIR =CNN NIR (x NIR );

[0136] f HSI =CNN HSI (x HSI );

[0137] Among them, x NIR and x HSI They are near-infrared images and hyperspectral image data after dimensionality reduction; CNN NIR and CNN HSI are the convolutional neural networks pre-trained for the corresponding images; f NIR and f HSI are the extracted near-infrared features and hyperspectral features respectively; W NIR and W HSI is the weight;

[0138] Using the joint feature stacking strategy, f NIR and f HSI Generate the fusion feature vector f by weighting u :

[0139] ;

[0140] Adding position encoding PE to the fusion feature allows the model to take advantage of the sequence order:

[0141] ;

[0142] Among them, pos refers to the position, is the dimension index, d model is the dimension of the feature.

[0143] In this embodiment, the multimodal detection model is as follows:

[0144] A multimodal detection model is built based on the improved R-CNN architecture, and a custom module is used to introduce channel attention mechanism and spatial attention mechanism to highlight the important features of defects;

[0145] Channel attention emphasizes the important channels in the feature map and generates channel attention M through global average pooling and global maximum pooling c (X):

[0146] M c (X) = σ s (W 1 (δ(W 0 (F a (X)))))+W 1 (δ(W 0 (F m (X))));

[0147] Among them, σ s is the Sigmoid activation function; δ is the nonlinear activation function; W 0 and W 1 is the weight matrix;

[0148] F a (X), F m (X) are the global average pooling and global maximum pooling operations of the input feature map X respectively;

[0149] Spatial AttentionM s (X), used to emphasize the area with significant features in the feature map:

[0150] M s (X) = σ s (f 7×7 ([F a (X);F m (X)]));

[0151] Among them, f 7×7 Indicates that a 7×7 convolution kernel is used for convolution operation;

[0152] Simultaneously learn the classification task and bounding box regression task;

[0153] Combine the output of the attention module to generate category probabilities for defect identification and bounding box predictions for precise positioning;

[0154] The loss function L includes the classification loss using the cross entropy loss L cls , regression loss uses smooth L1 loss L reg :

[0155] ;

[0156] in: is the classification loss, which is used to evaluate the predicted result p and the true label The difference between is the regression loss used to evaluate the predicted bounding box parameters With the real parameters The difference between is a hyperparameter; a represents the ath sample, Represents a binary switch variable that determines whether the a-th candidate box needs to calculate the regression loss of the bounding box.

[0157] In this embodiment, the R-CNN architecture is improved, including:

[0158] The region proposal network RPN predicts the objectness score and bounding box regression of the anchor box through a sliding window;

[0159] RoI pooling layer converts feature maps of candidate regions of different sizes into fixed sizes; RoI feature=pool(F,RoIs)

[0160] Where F is the input feature map, RoIs are candidate boxes generated by RPN, and RoI pooling adjusts the mapping according to the size of each RoI to make the output RoI feature size fixed;

[0161] A deep feature extraction network is used to perform a classification and regression joint training strategy to improve the detection accuracy, and fine classification and precise regression of the bounding box are performed for each RoI.

[0162] In this embodiment, S6 is specifically:

[0163] Assign a confidence score to each test result, set a confidence threshold, and filter out test results below the threshold; use a secondary classifier to reclassify the preliminary test results to further distinguish between real targets and falsely detected targets; combine scene knowledge or location restrictions (such as defects will not appear in a certain type of area) to filter out unreasonable test results and obtain the final test results;

[0164] According to the final test results, draw a box on the image to mark the defect location, mark the test type result and attach the test confidence level;

[0165] Perform pixel counting on the merged area to estimate the number of pixels occupied by defects, and create an automated module to generate a test report in a standard format, including the number, location, type, area and corresponding recommendations of defects.

[0166] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0167] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0168] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0170] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. An intelligent method for identifying defects on printed fabrics based on multimodal data, characterized in that: The following steps are involved: S1: Use a high-resolution RGB camera to obtain a global image of the printed cloth surface and perform preprocessing; S2: Based on the image recognition model, potential defect candidate areas are identified, and non-maximum suppression is combined to reduce overlapping defect candidate areas; S3: Acquire an image of the defect candidate area through a near-infrared camera device, and acquire a hyperspectral image through hyperspectral imaging; S4: extracting features from the near infrared image and hyperspectral image of the defect candidate area and performing feature fusion; S5: Based on the fused features and the multimodal detection model, fine-grained defect recognition and positioning are performed on the defect candidate areas, and an attention mechanism is added to improve the ability to capture complex features of defects; S6: Through post-processing steps, false positives are removed and adjacent areas are merged to improve the accuracy of defect detection. Finally, a test report is generated, marking the location, type and possible area of ​​the defect, so that the operator can make adjustments to the subsequent process; The image recognition model is built based on convolutional feature extraction network, region proposal network and Faster R-CNN, as follows: The convolutional feature extraction network uses the ResNet network to extract rich feature representations on the image; The region proposal network RPN generates a set of candidate regions on the feature map through a sliding window. When each window slides to a feature point, multiple candidate region anchors of different scales and aspect ratios are generated; RPN uses a binary classification model to determine whether each anchor contains a target and a specific defect category; The precise boundary of the candidate region is further adjusted through regression, and the loss function is: ; in, Candidate area The probability value of being predicted as a certain category; Candidate area The true category label of is the predicted bounding box regression parameter; is the true bounding box regression parameter; is the normalization factor of the classification loss, is the normalization factor of the regression loss, is the loss weight coefficient, is the classification loss function; is the regression loss function: Fast R-CNN extracts RoI features, and uses RoI Pooling to crop fixed-size features from the convolutional feature map based on the candidate regions generated by RPN. It also uses a fully connected layer to perform category discrimination and bounding box refinement on each candidate feature to obtain defect candidate regions. The multimodal detection model is specifically as follows: A multimodal detection model is built based on the improved R-CNN architecture, and a custom module is used to introduce channel attention mechanism and spatial attention mechanism to highlight the important features of defects; Channel attention emphasizes the important channels in the feature map and generates channel attention M through global average pooling and global maximum pooling c (X): M c (X)=σ s (W1(δ(W0(F a (X)))))+W1(δ(W0(F m (X)))); Among them, σ s is the Sigmoid activation function; δ is the nonlinear activation function; W0 and W1 are weight matrices; F a (X), F m (X) are the global average pooling and global maximum pooling operations of the input feature map X respectively; Spatial AttentionM s (X), used to emphasize the area with significant features in the feature map: M s (X)=σ s (f 7×7 ([F a (X);F m (X)])); Among them, f 7×7 Indicates that a 7×7 convolution kernel is used for convolution operation; Simultaneously learn the classification task and bounding box regression task; Combined with the output of the attention module, it generates category probabilities for defect recognition and bounding box predictions for precise positioning; The loss function L includes the classification loss using the cross entropy loss L cls , regression loss uses smooth L1 loss L reg : ; in: is the classification loss, which is used to evaluate the predicted result p and the true label The difference between is the regression loss used to evaluate the predicted bounding box parameters With the real parameters The difference between is a hyperparameter; a represents the ath sample, Represents a binary switch variable that determines whether the a-th candidate box needs to calculate the regression loss of the bounding box.

2. The method for intelligently identifying printed cloth surface defects based on multimodal data according to claim 1, characterized in that: The pre-processing is specifically as follows: Through the automatic white balance algorithm, the gains of the red, green and blue channels are adjusted in real time to ensure the authenticity and consistency of the colors; Use Gaussian filter to remove image noise, enhance image signal-to-noise ratio, and use Fourier transform to remove periodic noise; The image contrast is enhanced by histogram equalization, making defects easier to identify, and the gamma value is adjusted to optimize the overall brightness of the image, so as to better show the texture and color details of the fabric in the unsaturated color spectrum.

3. The method for intelligently identifying printed cloth surface defects based on multimodal data according to claim 2, characterized in that: The Gaussian filter is used to remove image noise, enhance the image signal-to-noise ratio, and the Fourier transform is used to remove periodic noise, as follows: The Gaussian filter convolves the image with a Gaussian kernel: ; Among them, 2g+1 is the size of the Gaussian kernel, I(x,y) represents the pixel value of the input image, I′(x,y) represents the pixel value of the output image; (x,y) is the position of the pixel in the image; (i,j) is the relative offset coordinate of the Gaussian kernel; represents Gaussian kernel; Fourier transform converts the image processed by Gaussian filtering from the spatial domain to the frequency domain, identifies and removes noise of specific frequencies, and performs a two-dimensional Fourier transform on the input image I′(x,y) to obtain the frequency domain representation F(u,v); ; Where M and N represent the width and height of the input image respectively; u and v are the indices of the two-dimensional frequency; is an imaginary unit; Identify and suppress high frequencies or specific frequency bands that represent periodic noise, and set the corresponding noise frequencies to zero in the Fourier plane: F′(u,v)=F(u,v)×H(u,v) Where H(u,v) is the frequency domain filter; Perform a two-dimensional inverse Fourier transform on the processed frequency domain image F′(u,v) and restore it to the spatial domain: 。 4. The method for intelligently identifying printed cloth surface defects based on multimodal data according to claim 1, characterized in that: The method of combining non-maximum suppression to reduce overlapping defect candidate regions is as follows: Arrange the bounding boxes of all detection results in descending order according to their category confidence. For a bounding box set {B k }, sorted in descending order of confidence: L(B)={B1,B2,...,B k ,...,B n ∣ score(B k )≥score(B k+1 )}; From the sorted candidate box set L(B), select the box B with the highest confidence in each cycle max ; Calculate overlap IOU: ; For all other boxes B k , if IOU(B k ,B max )>θ, θ is the preset threshold, then adjust the integral score : ; in, To control the parameters; Boxes with scores below a preset threshold are removed and the update calculation is repeated.

5. The method for intelligently identifying printed fabric surface defects based on multimodal data according to claim 1, characterized in that: The S4 is specifically: Standardize the input near-infrared image and hyperspectral image data, including distribution normalization for each band; The hyperspectral image data is subjected to dimensionality reduction through principal component analysis; CNN is applied to near-infrared images and hyperspectral image data after dimensionality reduction to extract basic feature representations; f NIR =CNN NIR (x NIR ); f HSI =CNN HSI (x HSI ); Among them, x NIR and x HSI They are near-infrared images and hyperspectral image data after dimensionality reduction; CNN NIR and CNN HSI are the convolutional neural networks pre-trained for the corresponding images; f NIR and f HSI are the extracted near-infrared features and hyperspectral features respectively; W NIR and W HSI is the weight; Using the joint feature stacking strategy, f NIR and f HSI Generate the fusion feature vector f by weighting u : I ; Adding position encoding PE to the fusion feature allows the model to take advantage of the sequence order: ; Among them, pos refers to the position, is the dimension index, d model is the dimension of the feature.

6. The method for intelligently identifying printed cloth surface defects based on multimodal data according to claim 1, characterized in that: The improved R-CNN architecture is as follows: The region proposal network RPN predicts the objectness score and bounding box regression of the anchor box through a sliding window; RoI pooling layer converts feature maps of candidate regions of different sizes into fixed sizes; RoI feature=pool(F,RoIs) Where F is the input feature map, RoIs are candidate boxes generated by RPN, and RoI pooling adjusts the mapping according to the size of each RoI to make the output RoI feature size fixed; A deep feature extraction network is used to perform a classification and regression joint training strategy to improve the detection accuracy, and fine classification and precise regression of the bounding box are performed for each RoI.

7. The method for intelligently identifying printed cloth surface defects based on multimodal data according to claim 1, characterized in that: The S6 is specifically: Assign a confidence score to each detection result, set a confidence threshold, and filter out detection results below the threshold; use a secondary classifier to reclassify the preliminary detection results to distinguish between real targets and falsely detected targets; Combine scene knowledge or location restrictions to filter out unreasonable detection results and obtain the final detection results; According to the final test results, draw a box on the image to mark the defect location, mark the test type result and attach the test confidence level; Perform pixel counting on the merged area to estimate the number of pixels occupied by defects, and create an automated module to generate a test report in a standard format, including the number, location, type, area and corresponding recommendations of defects.

Citation Information

Patent Citations

  • Fabric defect detection method based on hybrid expansion convolution and multi-scale feature fusion

    CN114067170A

  • Defect detection method based on improved Faster R-CNN

    CN114078106A

Cited By

  • Intelligent printing processing method and system based on image recognition

    CN121481974A