Fabric defect detection method based on spatial-temporal feature fusion enhancement
By introducing spatiotemporal feature fusion and multi-scale multi-frame consistency time convolution filtering technology in fabric defect detection, the problem of accuracy and robustness of defect detection in high-speed moving fabric scenes is solved, efficient and accurate defect detection is achieved, and the efficiency and quality of textile production is improved.
Patent Information
- Application Number
- CN202510290655.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In the high-speed moving fabric scene, the single-frame image information is limited, resulting in a significant reduction in the accuracy and robustness of fabric defect detection, and the dynamic information of the video stream is not fully utilized.
The fabric defect detection method based on spatial and temporal feature fusion enhancement is adopted. By obtaining fabric moving image sequence data, inputting the spatial and temporal feature fusion defect detection model, fusing spatial and temporal features, and extracting multi-scale multi-frame consistent time convolution filtering features to achieve defect detection.
It improves the accuracy and robustness of defect detection in high-speed moving fabric scenarios, reduces labor costs, reduces defect product rates, and improves production efficiency and product quality.
Smart Images

Figure CN120163802A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a fabric defect detection method based on video stream data. It is applicable to fabric quality detection and defect positioning in the textile industry, and is particularly suitable for achieving high-precision defect detection in the environment of fabrics moving at high speed. Background Art
[0002] In the modern textile industry, fabric defect detection has always been one of the key issues in development. The existence of defects will affect the aesthetics and quality of fabrics. Early detection and treatment of defects are crucial for textile production.
[0003] Fabric defect detection is an important part of quality control in the textile industry. The manual detection method relies on the human eye to observe and judge whether there are defects on the fabric surface. This method is simple to operate, but has problems such as low detection efficiency, low accuracy, and high cost. Especially when dealing with large-scale high-speed production tasks, its performance is particularly insufficient. The fabric inspection machine is a machine device used to detect fabric defects. It is mainly applied in the textile production line to check fabric surface defects such as holes, poor stitching, etc. to ensure the quality of textiles. The fabric inspection machine adopts high-speed imaging and image processing technologies, and can quickly and accurately identify the defects on the fabric, replacing the traditional manual detection method, greatly improving the detection efficiency and accuracy. At the same time, through the fabric inspection machine, for fabrics with complex fabric structures or delicate textures, they can be better detected and judged, effectively avoiding problems such as missed detection and misjudgment. The fabric inspection machine is one of the important devices in modern textile production, and is of great significance for ensuring textile quality, improving production efficiency, and reducing costs. In the prior art, Patent 1 with the patent publication number CN103456021B performs defect detection on a single-frame image based on morphological analysis. Patent 2 with the patent publication number CN113433137 uses the YOLO network to detect defects on the surface of a single-frame fabric; Patent 3 with the patent publication number CN109934802B extracts single-frame image defects by combining Fourier transform and image morphology. However, the disadvantages of the above prior art are that they are based on static images and use feature extraction, classification algorithms, or deep learning models to achieve automatic detection of fabric defects, without making full use of the dynamic information of the video stream. Although these methods have high accuracy in low-speed or static fabric scenarios, for fabrics moving at high speed, due to the limited information of a single-frame image, the detection accuracy and robustness are significantly reduced. In addition, traditional deep learning methods mainly focus on single-frame static features and ignore the time series information between consecutive frames.
[0004] In view of the above technical problems, no effective solution has been proposed yet. Summary of the Invention
[0005] An embodiment of the present invention provides a fabric defect detection method based on spatio-temporal feature fusion enhancement, which can greatly save labor costs, reduce the defective product rate, improve production efficiency and product quality, and bring economic benefits to textile enterprises.
[0006] The present invention achieves this purpose through the following technical solutions:
[0007] A fabric defect detection method based on spatio-temporal feature fusion enhancement, comprising the following steps:
[0008] Obtain the fabric moving image sequence data on at least one fabric inspection machine;
[0009] Input the fabric moving image sequence data into a spatio-temporal feature fusion defect detection model to obtain the defect detection result of the fabric moving image sequence data output by the spatio-temporal feature fusion defect detection model, where the defect detection result includes the defect type and the fabric physical position information corresponding to the defect;
[0010] The spatio-temporal feature fusion defect detection model includes a fabric image spatial alignment module, a fabric image grid division module, a pre-trained image encoder module, a defect feature enhancement module, a multi-scale multi-frame consistency temporal convolution filtering module, and a classification head; the fabric image spatial alignment module spatially aligns the fabric moving image sequence along the fabric moving direction, and each frame of image obtains a regional segment corresponding to the same fabric physical position; the fabric image grid division module divides the regional segment in each frame of the spatially aligned image sequence into several grid segments in the same way along the direction perpendicular to the fabric movement, takes out the grid segments corresponding to the same fabric physical position in each frame of image, and generates a grid segment image sequence of each fabric physical position at different times; the pre-trained image encoder module is used to extract features from the grid segment image sequence to generate a grid segment feature vector sequence; the defect feature enhancement module enhances the feature of the grid segment feature vector sequence by introducing a learnable background feature vector in the feature space to obtain a grid segment enhanced feature vector sequence; the multi-scale multi-frame consistency temporal convolution filtering module processes the grid segment enhanced feature vector sequence using a multi-scale convolution strategy, performs parallel convolution in time windows of different durations, so as to better model defects of different durations; the classification head sequentially processes through feature flattening, a fully connected layer, and a softmax activation function to obtain a classification result. For the grid segment classified as a defect, its corresponding fabric physical position information is obtained according to the counter.
[0011] Further, the spatio-temporal feature fusion defect detection model is obtained by the following steps:
[0012] S1: Construct a defect detection dataset for the fabric image sequence, preprocess the image sequence data and divide it into a training set, a validation set and a test set;
[0013] S2: Input the fabric image sequence data into the fabric image spatial alignment module to obtain a sequence of regional segment images, where each regional segment image corresponds to the sampling of the same fabric physical position at different times;
[0014] S3: Input the sequence of regional segment images into the fabric image grid division module to obtain multiple groups of grid segment image sequences, where each group of grid segment image sequences corresponds to the sampling of the same fabric physical position at different times;
[0015] S4: Input each group of grid segment image sequences into the pre-trained image encoder module respectively to obtain a sequence of grid segment feature vectors;
[0016] S5: Input each group of grid segment feature vector sequences into the defect feature enhancement module respectively to obtain a sequence of grid segment enhanced feature vectors;
[0017] S6: Input each group of grid segment enhanced feature vector sequences into the multi-scale multi-frame consistency temporal convolutional filtering module respectively to obtain a sequence of grid segment filtered feature vectors;
[0018] S7: Input the sequence of grid segment filtered feature vectors into the classification head to obtain the grid segment classification result;
[0019] S8: Calculate the loss according to the model output result and the manual annotation result, perform gradient backpropagation, and update the model parameters;
[0020] S9: Repeat steps S2 to S8, iteratively train until the model converges, and the training stage ends;
[0021] S10: Use the validation set to select the optimal hyperparameters for the model;
[0022] S11: Evaluate the generalization performance of the model using the test set.
[0023] Furthermore, in the step S1, the system acquires the image sequence on the fabric surface in the form of continuous frames through a camera fixedly installed on the fabric inspection machine. represents the fabric image captured by the camera at time t. H, W, and C represent the height, width, and number of channels of the image respectively. It is set that the direction from left to right of the collected fabric image is the x-axis direction, the direction from top to bottom is the y-axis direction, the upper left corner coordinates of the image are (0,0), and the lower right corner coordinates are (W-1,H-1);
[0024] Specifically, the image acquired by the camera is an RGB image, so the number of channels C = 3;
[0025] Specifically, the fabric moves along the x-axis at a known constant speed v (pixel / frame), and the sampling time interval between two adjacent frames of images is Δt (seconds);
[0026] Specifically, the fabric image sequence corresponding to any time T0 is Contains T frames of images acquired continuously.
[0027] During manual labeling, only the first frame of each fabric image sequence needs to be labeled;
[0028] Further, in step S2, for the fabric image sequence In the fabric image acquired at time T0+kΔt (k∈{1,2,...,T}), a rectangular area with upper left corner coordinates ((k-1)v,0) and lower right corner coordinates ((k-1)v+w-1,H-1) is intercepted to form the kth area segment, forming an area segment image sequence consisting of T area segments.
[0029] Further, in step S3, the region segment image sequence composed of T region segments is The grid is divided into M grid segments along the y-axis direction, and the height of each grid segment is h = H / M. In the fabric image acquired at time T0+kΔt (k∈{1,2,...,T}), the m∈{1,2,...,M}th grid segment The coordinates of the upper left corner are ((k-1)v, (m-1)h), and the coordinates of the lower right corner are ((k-1)v+w-1, (m-1)h+h-1). Get the region fragment image sequence The mth grid segment of each region segment in forms a grid segment image sequence consisting of T grid segments.
[0030] Further, in step S4, the pre-trained image encoder module can be expressed as:
[0031]
[0032] Where θ represents the parameters of the pre-trained network and d is the output feature dimension. The pre-trained image encoder module is used to transform each grid segment image Convert to a d-dimensional feature vector:
[0033]
[0034] The feature vectors of T moments are collected to obtain the grid segment feature vector sequence:
[0035]
[0036] Further, in the step S5, since industrial fabrics often have a wide range of repetitive textures and a relatively stable normal state, a learnable background feature vector is introduced in the feature space for benchmark modeling of common textures in the "flawless" case. This vector can be updated during the training process to make it as close as possible to the most common normal background distribution in industrial production.
[0037] Specifically, the calculation method of the grid segment enhanced feature vector is as follows:
[0038]
[0039] where is the feature vector of the enhanced flaw information at time T0 + kΔt.
[0040] Specifically, if there are no flaws in the input grid segment , then should be close to z template , and at this time, each element in is close to 0; otherwise, if there are flaws, the difference between them increases, and at this time, significant non-zero components will appear in .
[0041] Specifically, the grid segment enhanced feature vectors at T moments are pooled to obtain a grid segment enhanced feature vector sequence:
[0042]
[0043] Further, in the step S6, multi-frame consistency filtering is performed on the grid segment enhanced feature vector sequence in the time dimension to strengthen the flaw signals that appear across frames by aggregating the convolution responses at different time scales, and suppress the noise spikes that only appear in a single frame or a very short period, so as to obtain a more discriminative flaw feature representation.
[0044] Specifically, for the grid segment enhanced feature vector , B parallel convolution calculation branches Conv1D (b) with different convolution kernel sizes A (b) (b ∈ {1, 2,..., B}) are set, and 1D convolution calculations are respectively performed in the time dimension:
[0045]
[0046] where C′ represents the number of channels of each convolution kernel.
[0047] Specifically, aggregate the calculation results of the parallel convolution calculation branches to obtain the network segment filtering feature vector:
[0048]
[0049] where α b represents the weight of the trainable convolution calculation branch, and σ(·) represents the activation function.
[0050] Specifically, pool the network segment filtering feature vectors at T moments to obtain the network segment filtering feature vector sequence:
[0051]
[0052] Furthermore, in step S7, input the network segment filtering feature vector sequence into the fully connected classifier to obtain the classification result.
[0053] Specifically, flatten the network segment filtering feature vector sequence into a one-dimensional vector Suppose there are Q - 1 types of defects, plus the non-defective category, for a total of Q categories. Use a fully connected layer to process and use the softmax activation to obtain the probabilities of being classified into each category:
[0054]
[0055] where, and represent the weight and bias of the fully connected layer respectively; represents the prediction result vector, represents the network segment image sequence the probability of being classified into category q. Take the category with the highest probability as the model classification result.
[0056] Furthermore, in step S8, the model is trained using the classification loss and the consistency loss.
[0057] Specifically, for the fabric image sequence corresponding to any moment T0 its classification loss can be expressed as:
[0058]
[0059] where, is a one-hot vector. If has a value of 1, it means that the category label of the network segment image sequence is q, represents the time index of all samples in the training set, Indicates the size of the training set.
[0060] Specifically, a consistency loss is introduced to constrain z on the sequence of defect-free grid segment images with template the similarity of :
[0061]
[0062] Wherein, represents the data set formed by the sequence of defect-free grid segment images, represents the size of
[0063] Specifically, the classification loss and the consistency loss are weighted and combined as the total loss:
[0064] L = L cls + λL cons
[0065] Wherein, λ represents the loss balance coefficient.
[0066] Compared with the prior art, the core innovation of the present invention lies in introducing the fusion of time series analysis and multi-frame information of video stream data. By extracting features and modeling hidden states for consecutive frames in the same spatial region, the present invention additionally fuses the dynamic evolution features in the time dimension on the premise of unchanged spatial resolution, enabling the detection model to learn the temporal law of defect features from cross-frame information. In addition, the proposed multi-scale multi-frame consistency time convolution filtering module can highlight the defect features that continuously appear in consecutive frames while smoothing out occasional random noise that may be caused by environmental factors. Compared with the prior art, the present invention has stronger robustness and generalization ability for defect detection under high-speed fabric movement. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0068] Figure 1 is a flowchart implemented according to the present invention;
[0069] Figure 2 is a structural diagram of the spatio-temporal feature fusion defect detection model designed according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0071] The present invention discloses a fabric defect detection method based on spatio-temporal feature fusion enhancement, including the following:
[0072] Obtain the fabric moving image sequence data on at least one fabric inspection machine;
[0073] Input the fabric moving image sequence data into the spatio-temporal feature fusion defect detection model to obtain the defect detection result of the fabric moving image sequence data output by the spatio-temporal feature fusion defect detection model. The defect detection result includes the defect type and the fabric physical position information corresponding to the defect.
[0074] The spatio-temporal feature fusion defect detection model includes a fabric image spatial alignment module, a fabric image grid division module, a pre-trained image encoder module, a defect feature enhancement module, a multi-scale multi-frame consistent temporal convolution filtering module, and a classification head; the fabric image spatial alignment module spatially aligns the fabric moving image sequence along the fabric moving direction, and each frame of image obtains a region segment corresponding to the same fabric physical position; the fabric image grid division module divides the region segment in each frame of the spatially aligned image sequence into several grid segments in the same way along the direction perpendicular to the fabric movement, takes out the grid segments corresponding to the same fabric physical position in each frame of image, and generates a grid segment image sequence of each fabric physical position at different times; the pre-trained image encoder module is used to extract features from the grid segment image sequence to generate a grid segment feature vector sequence; the defect feature enhancement module enhances the grid segment feature vector sequence by introducing a learnable background feature vector in the feature space to obtain a grid segment enhanced feature vector sequence; the multi-scale multi-frame consistent temporal convolution filtering module processes the grid segment enhanced feature vector sequence by using a multi-scale convolution strategy, and performs parallel convolution in time windows of different durations, so as to better model defects of different durations; the classification head sequentially processes through feature flattening, a fully connected layer, and a softmax activation function to obtain a classification result. For the grid segments classified as defects, their corresponding fabric physical position information is obtained according to the counter.
[0075] The overall inference process of the spatio-temporal feature fusion defect detection model is as follows: Input the fabric image sequence data into the fabric image spatial alignment module to obtain a sequence of regional fragment images; input the sequence of regional fragment images into the fabric image grid division module to obtain multiple groups of grid fragment image sequences; input each group of grid fragment image sequences into the pre-trained image encoder module respectively to obtain a sequence of grid fragment feature vectors; input each group of grid fragment feature vectors into the defect feature enhancement module respectively to obtain a sequence of grid fragment enhanced feature vectors; input each group of grid fragment enhanced feature vectors into the multi-scale multi-frame consistency temporal convolutional filtering module respectively to obtain a sequence of grid fragment filtered feature vectors; input the sequence of grid fragment filtered feature vectors into the classification head to obtain the grid fragment classification result.
[0076] The training process of the spatio-temporal feature fusion defect detection model is as Figure 1 、 Figure 2 shown, and the specific steps are as follows:
[0077] S1: Construct a fabric image sequence defect detection dataset, preprocess the image sequence data, and divide it into a training set, a validation set, and a test set.
[0078] Specifically, first use a camera fixedly installed on the fabric inspection machine to obtain an image sequence of the fabric surface in the form of continuous frames where Δt represents the sampling time interval between two adjacent frames of images, represents the fabric image captured by the camera at time T0 + kΔt, and H, W, and C respectively represent the height, width, and number of channels of the image; then process the first frame of fabric image in the image sequence to obtain a sequence of grid fragment images of its leftmost region. Set the left-to-right direction of the collected fabric image as the x-axis direction and the top-to-bottom direction as the y-axis direction. The upper-left and lower-right coordinates of the region corresponding to each grid fragment in the image are (0, (m - 1)h) and (w - 1, (m - 1)h + h - 1) (m = 1, 2,..., M) respectively. Where w represents the width of each grid fragment, h represents the height of each grid fragment, and M = H / h represents the number of grid fragments; then, manually annotate each grid fragment, and the annotation result is the defect category or no defect; finally, perform the division of the dataset. A total of R groups of image sequences are collected The grid fragments obtained from a certain proportion (such as 60%) of the earliest collected image sequences are used as the training set, the grid fragments obtained from a certain proportion (such as 20%) of the image sequences collected in the middle period are used as the validation set, and the grid fragments obtained from a certain proportion (such as 20%) of the image sequences collected at the latest are used as the test set.
[0079] S2: Input the fabric image sequence data into the fabric image spatial alignment module to obtain a sequence of regional fragment images, where each regional fragment image corresponds to the sampling of the same fabric physical position at different times.
[0080] Specifically, for the fabric image sequence In the fabric image obtained at the moment of T0 + kΔt (k ∈ {1, 2,..., T}), intercept a rectangular area with the upper left coordinate ((k - 1)v, 0) and the lower right coordinate ((k - 1)v + w - 1, H - 1) to form the k-th regional fragment, and form a sequence of regional fragment images composed of T regional fragments.
[0081] S3: Input the sequence of regional fragment images into the fabric image grid division module to obtain multiple groups of grid fragment image sequences, where each group of grid fragment image sequences corresponds to the sampling of the same fabric physical position at different times.
[0082] Specifically, for the sequence of regional fragment images composed of T regional fragments Perform grid division along the y-axis direction, divided into M grid fragments in total, and the height of each grid fragment is h = H / M. In the fabric image obtained at the moment of T0 + kΔt (k ∈ {1, 2,..., T}), the m-th grid fragment The upper left coordinate is ((k - 1)v, (m - 1)h), and the lower right coordinate is ((k - 1)v + w - 1, (m - 1)h + h - 1). Obtain the m-th grid fragment of each regional fragment in the sequence of regional fragment images to form a sequence of grid fragment images composed of T grid fragments.
[0083] S4: Input each group of grid fragment image sequences into the pre-trained image encoder module respectively to obtain a sequence of grid fragment feature vectors.
[0084] Specifically, the pre-trained image encoder module can be expressed as:
[0085]
[0086] where θ represents the parameters of the pre-trained network, and d is the output feature dimension. The pre-trained image encoder module is used to convert each grid fragment image into a d-dimensional feature vector:
[0087]
[0088] Pool the feature vectors at T moments to obtain a sequence of grid fragment feature vectors:
[0089]
[0090] S5: Input each group of grid segment feature vector sequences into the defect feature enhancement module respectively to obtain grid segment enhanced feature vector sequences.
[0091] Specifically, the calculation method of the grid segment enhanced feature vector is as follows:
[0092]
[0093] Among them, is the feature vector of the enhanced defect information at time T0 + kΔt.
[0094] If there is no defect in the input grid segment , then should be close to z template . At this time, each element in is close to 0; otherwise, if there is a defect, the difference between the two increases, and at this time there will be significant non-zero components in.
[0095] Pool the grid segment enhanced feature vectors at T time instants to obtain a grid segment enhanced feature vector sequence:
[0096]
[0097] S6: Input each group of grid segment enhanced feature vector sequences into the multi-scale multi-frame consistency temporal convolutional filtering module respectively to obtain grid segment filtered feature vector sequences.
[0098] Specifically, for the grid segment enhanced feature vector set up B parallel convolutional calculation branches Conv1D (b) with different convolutional kernel sizes A (b) (b ∈ {1, 2,..., B}), and perform 1D convolutional calculations respectively in the time dimension:
[0099]
[0100] Among them, C′ represents the number of channels of each convolutional kernel.
[0101] Aggregate the calculation results of the parallel convolutional calculation branches to obtain the network segment filtered feature vector:
[0102]
[0103] Among them, α b represents the weight of the trainable convolutional calculation branch, and σ(·) represents the activation function.
[0104] Pool the filtered feature vectors of the grid segments at T moments to obtain a sequence of filtered feature vectors of the grid segments:
[0105]
[0106] S7: Input the sequence of filtered feature vectors of the grid segments into the classification head to obtain the classification result of the grid segments.
[0107] Specifically, flatten the sequence of filtered feature vectors of the grid segments into a one-dimensional vector Suppose there are Q - 1 types of defects, plus the non-defective category, for a total of Q categories. Use a fully connected layer to process and use the softmax activation to obtain the probabilities of being classified into each category:
[0108]
[0109] Among them, and represent the weights and biases of the fully connected layer respectively; represents the prediction result vector, represents the sequence of grid segment images The probability of being classified into category q. Take the category with the highest probability as the model classification result.
[0110] S8: According to the model output result and the manually annotated result, calculate the loss, perform gradient backpropagation, and update the model parameters.
[0111] Specifically, the model training uses classification loss and consistency loss. For the fabric image sequence corresponding to any moment T0 Its classification loss can be expressed as:
[0112]
[0113] Among them, is a one-hot vector. If has a value of 1, it means that the category label of the grid segment image sequence is q, represents the moment index of all samples in the training set, represents the size of the training set.
[0114] The consistency loss can be expressed as:
[0115]
[0116] Among them, represents the data set formed by the non-defective grid segment image sequences, represents Dimensions.
[0117] The classification loss and the consistency loss are weighted and combined as the total loss:
[0118] L = L cls + λL cons
[0119] where λ represents the loss balance coefficient.
[0120] S9: Repeat steps S2 to S8, iteratively train until the model converges, and the training phase ends.
[0121] S10: Use the validation set to select the optimal hyperparameters for the model.
[0122] Specifically, different hyperparameters are set for the model to obtain different model instances; for each model instance, train on the training set according to steps S2 to S9 and perform inference on the validation set, and calculate the performance metrics of each model instance on the validation set; compare the performance metrics of each model instance on the validation set, and select the model instance with the optimal performance metrics as the finally used model.
[0123] S11: Evaluate the generalization performance of the model using the test set.
[0124] Specifically, for the finally used model selected through the validation set, perform model inference on the test set and calculate the performance metrics of the model on the test set.
[0125] The above has described the present invention in detail through embodiments, but the content described is only an exemplary embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. All those who utilize the technical solutions described in the present invention, or those skilled in the art inspired by the technical solutions of the present invention, within the essence and protection scope of the present invention, design similar technical solutions to achieve the above technical effects, or make equivalent changes and improvements to the application scope, etc., should still fall within the scope covered by the patent of the present invention.
Claims
1. A fabric defect detection method based on spatiotemporal feature fusion enhancement, characterized in that: The following steps are involved: Acquiring fabric moving image sequence data on at least one fabric inspection machine; Inputting the fabric moving image sequence data into a spatiotemporal feature fusion defect detection model to obtain a defect detection result of the fabric moving image sequence data output by the spatiotemporal feature fusion defect detection model, wherein the defect detection result includes defect type and fabric physical location information corresponding to the defect; The spatiotemporal feature fusion defect detection model includes a fabric image spatial alignment module, a fabric image grid division module, a pre-trained image encoder module, a defect feature enhancement module, a multi-scale multi-frame consistency temporal convolution filtering module, and a classification head; the fabric image spatial alignment module spatially aligns the fabric moving image sequence along the fabric moving direction, and each frame of the image obtains a region segment corresponding to the same fabric physical position; The fabric image grid division module divides the regional fragments in each frame of the spatially aligned image sequence into a number of grid fragments in the same manner and along a direction perpendicular to the movement of the fabric, and extracts the grid fragments corresponding to the same physical position of the fabric in each frame of the image to generate a grid fragment image sequence at each physical position of the fabric at different times; the pre-trained image encoder module is used to extract features from the grid fragment image sequence to generate a grid fragment feature vector sequence; the defect feature enhancement module enhances the features of the grid fragment feature vector sequence by introducing a learnable background feature vector in the feature space to obtain a grid fragment enhanced feature vector sequence; the multi-scale multi-frame consistent temporal convolution filtering module processes the grid fragment enhanced feature vector sequence using a multi-scale convolution strategy and performs parallel convolution on time windows of different lengths, so as to better model defects of different durations; the classification head obtains the classification result through feature flattening, full connection layer, and softmax activation function processing in sequence; for the grid fragment classified as a defect, the corresponding fabric physical position information is obtained according to the meter.
2. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 1, characterized in that: The spatiotemporal feature fusion defect detection model is obtained by the following steps: S1: Construct a fabric image sequence defect detection dataset, preprocess the image sequence data and divide it into training set, validation set and test set; S2: inputting the fabric image sequence data into the fabric image spatial alignment module to obtain a sequence of regional fragment images, wherein each regional fragment image corresponds to the sampling of the same fabric physical position at different times; S3: inputting the region segment image sequence into a fabric image grid division module to obtain multiple groups of grid segment image sequences, wherein each group of grid segment image sequences corresponds to sampling of the same fabric physical position at different times; S4: input each group of grid segment image sequences into the pre-trained image encoder module to obtain a grid segment feature vector sequence; S5: inputting each group of grid segment feature vector sequences into the defect feature enhancement module to obtain a grid segment enhanced feature vector sequence; S6: input each group of grid segment enhanced feature vector sequences into a multi-scale multi-frame consistent temporal convolution filtering module to obtain a grid segment filtered feature vector sequence; S7: input the mesh segment filtering feature vector sequence into the classification head to obtain the mesh segment classification result; S8: Calculate the loss, back propagate the gradient, and update the model parameters based on the model output results and manual annotation results; S9: Repeat steps S2 to S8, iterate the training until the model converges, and the training phase ends; S10: Use the validation set to select the optimal hyperparameters for the model; S11: Use the test set to evaluate the generalization performance of the model.
3. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 2 is characterized in that: In step S1, the system obtains an image sequence of the fabric surface in the form of continuous frames through a camera fixedly installed on the fabric inspection machine. It represents the fabric image taken by the camera at time t, and H, W, and C represent the height, width, and number of channels of the image, respectively. Set the direction of the collected fabric image from left to right as the x-axis direction, the direction from top to bottom as the y-axis direction, the coordinates of the upper left corner of the image as (0,0), and the coordinates of the lower right corner as (W-1,H-1); The image collected by the camera is an RGB image, so the number of channels C=3; the fabric moves along the x-axis at a known constant speed v (pixel / frame), and the sampling time interval between two adjacent frames of images is Δt (seconds); The fabric image sequence corresponding to any time T0 Contains T frames of images acquired continuously. During manual labeling, only the first frame of each fabric image sequence needs to be labeled.
4. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 2, characterized in that: In step S2, for the fabric image sequence In the fabric image acquired at time T0+kΔt (k∈{1,2,...,T}), a rectangular area with upper left corner coordinates ((k-1)v,0) and lower right corner coordinates ((k-1)v+w-1,H-1) is intercepted to form the kth area segment, forming an area segment image sequence consisting of T area segments. In step S3, a region segment image sequence consisting of T region segments is The grid is divided into M grid segments along the y-axis direction, and the height of each grid segment is h = H / M. In the fabric image acquired at time T0+kΔt (k∈{1,2,...,T}), the m∈{1,2,...,M}th grid segment The coordinates of the upper left corner are ((k-1)v, (m-1)h), and the coordinates of the lower right corner are ((k-1)v+w-1, (m-1)h+h-1). Get the region fragment image sequence The mth grid segment of each region segment in forms a grid segment image sequence consisting of T grid segments. In step S4, the pre-trained image encoder module can be expressed as: Where θ represents the parameters of the pre-trained network and d is the output feature dimension. The pre-trained image encoder module is used to transform each grid segment image Convert to a d-dimensional feature vector: The feature vectors of T moments are collected to obtain the grid segment feature vector sequence: In step S5, since industrial fabrics often have a large range of repeated textures and a relatively stable normal state, a learnable background feature vector is introduced into the feature space. Used to benchmark common textures in "defect-free" situations, this vector can be updated during training to make it as close as possible to the most common normal background distribution in industrial production.
5. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 4, characterized in that: The calculation method of the mesh segment enhanced feature vector is: in, That is, it is the feature vector of enhanced defect information at the time T0+kΔt; If you enter a mesh fragment If there are no defects in Should be with z template Close, now Each element in is close to 0; otherwise, if there is a defect, the difference between the two increases. There will be significant non-zero components in ; the grid segment enhanced feature vectors of T moments are collected to obtain the grid segment enhanced feature vector sequence: 。 6. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 2, characterized in that: In step S6, the feature vector sequence is enhanced for the mesh segment Perform multi-frame consistency filtering in the time dimension to enhance defect signals that appear across frames and suppress noise spikes that appear only in a single frame or a very short period of time by aggregating convolution responses at different time scales, thereby obtaining more discriminative defect feature representations; Enhanced feature vector for mesh fragments Set up B parallel convolutions with different kernel sizes A (b) The convolution calculation branch Conv1D (b) (b∈{1,2,...,B}), perform 1D convolution calculations in the time dimension: Among them, C′ represents the number of channels of each convolution kernel; Aggregate the calculation results of the parallel convolution calculation branches to obtain the network segment filtering feature vector: Among them, α b represents the trainable convolution calculation branch weight, and σ(·) represents the activation function; The grid segment filter feature vectors at T moments are collected to obtain a grid segment filter feature vector sequence: 。 7. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 2, characterized in that: In step S7, the grid segment filtering feature vector sequence Input the fully connected classifier to obtain the classification result; Filter the feature vector sequence of the mesh fragments Flatten into a one-dimensional vector Suppose there are Q-1 types of defects, plus the defect-free category, for a total of Q categories. Use a fully connected layer to Process it and use softmax activation to get the probability of classification into each category: in, and Represent the weight and bias of the fully connected layer respectively; represents the prediction result vector, Represents a grid fragment image sequence The probability of being classified as category q, taking the category with the highest probability As the model classification result.
8. The fabric defect detection method based on spatiotemporal feature fusion enhancement according to claim 2, characterized in that: In step S8, the model training uses classification loss and consistency loss; For any fabric image sequence corresponding to time T0 Its classification loss can be expressed as: in, is a one-hot vector if A value of 1 indicates a grid fragment image sequence The category label is q, Represents the time index of all samples in the training set, represents the size of the training set; introduces consistency loss to find the best fit in the sequence of defect-free mesh fragment images Upper constraint z template and Similarity: in, represents a data set formed by a sequence of images of defect-free mesh fragments, express Size; The classification loss and consistency loss are weighted together as the total loss: L=L cls +λL cons Where λ represents the loss balance coefficient.
Citation Information
Patent Citations
A cloth defect detection method based on morphological analysis
CN103456021B
A fabric defect detection method based on Fourier transform and image morphology
CN109934802B
Fabric flaw detection method based on light field camera depth information extraction
CN110349132A
Fabric defect detection method based on step-by-step identification strategy
CN113920098A
Method for detecting defects of plain yarn-dyed fabric
CN117495836A