High-speed movement fabric defect image generation method based on style migration network
Through the method based on the style migration network, high-fidelity high-speed moving fabric defect images are generated, which solves the problem of insufficient generalization ability of the target detection model in the high-speed operating environment in the prior art, and achieves efficient image generation and model generalization ability improvement.
Patent Information
- Application Number
- CN202510209820.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively generate high-speed moving fabric defect images, resulting in insufficient generalization capabilities of the target detection model in a high-speed operating environment and unable to meet industrial production needs.
Using a style migration network-based method, a multi-layer perceptron MLP and U-Net generator is used to combine Dino-ViT pre-trained feature extraction networks, and a feature reconstruction loss and style migration framework are used to generate high-fidelity high-speed moving fabric defect images.
It realizes the generation of high-quality high-speed moving fabric defect images on the basis of reducing the demand for hardware resources and reducing the number of images of high-speed moving fabric defects, which improves the generalization ability of the target detection model in a high-speed operating environment and meets industrial production needs.
Smart Images

Figure CN120219872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image generation method, in particular to a method for generating high-speed moving fabric defect images based on a style transfer network, belonging to the technical field of deep learning sample generation.
Background Art
[0002] With the continuous development of the textile industry, the production scale of cloth is constantly expanding, and at the same time, the quality requirements for cloth are getting higher and higher. Therefore, the defect detection of cloth has become an essential part of the production process. Traditional cloth defect detection relies on manual operation. However, the accuracy and efficiency of manual detection are often low, and the detection accuracy is usually between 50% and 70%.
[0003] To improve the detection accuracy and efficiency, object detection models have begun to be widely used in the field of cloth defect detection. However, the training of object detection models often requires a large number of cloth images with defects. In recent years, with the continuous improvement of the image acquisition system of the cloth inspection machine, the cloth inspector can immediately stop the machine and sample after discovering a defect, so as to obtain relatively rich static defect images. After training based on these images, the model shows excellent performance in defect detection rate under the low-speed movement state of the cloth. However, as the movement speed of the cloth increases, the defect detection rate drops severely and cannot meet the requirements of industrial production.
[0004] To improve the generalization of the model under high-speed operation, using high-speed moving fabric defect images to train the detection model is a solution with relatively low cost. However, limited by the observation ability of the cloth inspector and the hardware performance, it is extremely difficult to collect high-speed moving fabric defect images. Therefore, it is very necessary to design an efficient sample generation method that can effectively capture the characteristics of high-speed moving fabric defect images.
[0005] The existing fabric defect image generation network models are mainly variants based on the generative adversarial network (GAN) and autoencoders. However, the training of GAN requires too much hardware resources, and these methods often cannot generate large-area cloth defects with a large aspect ratio at one time. In addition, the training of these models all requires a large number of training samples, and it is difficult to collect high-speed moving fabric defect images, making it difficult to meet the quantity requirements of training samples.
[0006] Therefore, to solve the above problems, it is indeed necessary to provide an innovative method for generating high-speed moving fabric defect images based on a style transfer network to overcome the defects in the prior art.
Summary of the Invention
[0007] The object of the present invention is to provide a method for generating high-speed moving fabric defect images based on a style transfer network, which realizes the high-fidelity generation of high-speed moving fabric defect images on the basis of reducing the hardware resources required for training and reducing the dependence on the number of high-speed moving fabric defect images, improves the generalization ability of the target detection model in a high-speed working environment, injects new impetus into the industrial development, and promotes the intelligent upgrading of traditional industries.
[0008] To achieve the above object, the technical solution adopted by the present invention is: a method for generating high-speed moving fabric defect images based on a style transfer network, which includes the following process steps:
[0009] 1), Use an industrial camera to collect static fabric defect images and high-speed moving fabric defect images, take the static fabric defect images as source images, and the high-speed moving fabric defect images as target images to make a training set;
[0010] 2), Perform feature analysis and decoupling on the static and high-speed moving fabric defect images, and construct an MVU-Style-Transfer network model. The construction process is as follows: Use a multi-layer perceptron MLP, combined with a feature reconstruction loss to simulate the texture appearance imaging under high-speed movement of the fabric; Use U-Net as the generator, adopt a multi-source single-target style transfer framework, combined with the Dino-ViT pre-trained feature extraction network, and use the self-similarity metric based on the deepest layer key of the network and the L2 norm distance of the CLS token as the structure loss function and the visual appearance loss function to train the generator to obtain the MVU-Style-Transfer network.
[0011] 3), Use the training set to train the MVU-Style-Transfer network model to obtain a trained MVU-Style-Transfer network model;
[0012] 4), Input the static fabric defect image to be generated into the trained MVU-Style-Transfer network model to obtain the generation result of the high-speed moving fabric defect image.
[0013] The method for generating high-speed moving fabric defect images based on the style transfer network of the present invention is further: The step 1) is specifically:
[0014] 1-1), Use an industrial camera to collect several static fabric images with defects, adjust the speed of the fabric inspection machine to 60 m / min, and collect several high-speed moving fabric images with defects;
[0015] (1-2), perform data augmentation on the static fabric images to obtain enhanced static fabric images; expand all the collected static fabric images and the enhanced static fabric images into a static fabric image dataset, divide them by defect categories and use them as the source image set. Select a single image from the source image set, and use a single high-speed moving fabric image of the same defect category as the target image to construct a single-source single-target image set. Finally, construct a data combination of the same category of defects. The number of image sets in the data combination of each category of defects is not less than 1000, and finally a training set for each category is obtained.
[0016] The method for generating high-speed moving fabric defect images based on the style transfer network of the present invention is further as follows: in step 1-2), the data augmentation methods are flipping, rotating, adding noise, changing contrast, and changing brightness.
[0017] The method for generating high-speed moving fabric defect images based on the style transfer network of the present invention is further as follows: in step 2), specifically, the image is decoupled into texture motion appearance, visual appearance, and structure, and an MVU-Style-Transfer network model is constructed based on these three parts.
[0018] The method for generating high-speed moving fabric defect images based on the style transfer network of the present invention is further as follows: in step 2), the specific method for texture appearance imaging is: based on the input static fabric defect image, synthesize an image sequence that moves down 1 pixel value in sequence. The sequence length is equal to the input dimension of the multi-layer perceptron MLP. Randomly initialize a weight sequence of the same length, send the weight sequence into the multi-layer perceptron MLP, and after normalizing the weight sequence output by the multi-layer perceptron MLP, fuse it with the image sequence by weighting to obtain the synthesized intermediate image;
[0019] Among them, the input dimension of the MLP is calculated from the fabric movement speed and the camera exposure time:
[0020] N = TVk
[0021]
[0022] Among them, N represents the dimension of the input layer and the output layer, T represents the exposure time used by the camera, V represents the fabric movement speed, and k represents the ratio of the camera frame size to the fabric image size; I composite represents the intermediate image output by the MLP, α n , I n respectively represent the weight sequence of the input layer of the multi-layer perceptron MLP and the image sequence synthesized based on the input static fabric defect image. MLP(·) represents the weight sequence returned after fitting by the multi-layer perceptron, and softmax(·) represents SoftMax normalization of the weight sequence output by the multi-layer perceptron MLP.
[0023] The method for generating high-speed moving fabric defect images based on a style transfer network of the present invention further includes: in the step 2), the calculation formula of the feature reconstruction loss is as follows:
[0024]
[0025] where represents the feature reconstruction loss, specifically referring to the difference between I composite and I a in the j-th layer feature space of the pre-trained network φ. I composite and I a respectively represent the intermediate image output by the MLP and the high-speed moving fabric defect image. φ j (·) represents the feature mapping of the j-th layer of the VGG network, and ||·||2 represents the L2 norm distance between the feature mappings. C j , H j , and W j respectively represent the number of channels, height, and width of the j-th layer feature mapping.
[0026] The method for generating high-speed moving fabric defect images based on a style transfer network of the present invention further includes: in the step 2), the structure loss function and the visual appearance loss function satisfy the following formula:
[0027]
[0028] where L app represents the visual appearance loss function, represents the [CLS] token extracted from the L-th layer of the ViT model. I a and I o respectively represent the high-speed moving fabric defect image and the image generated by the generator. ||·||2 represents the L2 norm calculation; S L (I) ij represents the self-similarity between patch i and patch j of the image I in the L-th layer of the ViT model. cos-sim(·) represents the cosine similarity calculation, and respectively represent the keys of patch i and patch j in the L-th layer. L structure represents the structure loss function, and ||·|| F represents the F norm calculation. I s and I o respectively represent the static fabric defect image and the image generated by the generator.
[0029] The method for generating high-speed moving fabric defect images based on a style transfer network of the present invention is further as follows: In step 2), the multi-layer perceptron MLP includes an input layer, a hidden layer, and an output layer. The dimension of the input layer is calculated from the fabric movement speed and the camera exposure time. The hidden layer consists of three fully connected layers with 128 nodes each and the activation function Relu. The dimension of the output layer is the same as that of the input layer, and SoftMax is used to normalize the output.
[0030] The method for generating high-speed moving fabric defect images based on a style transfer network of the present invention is further as follows: The U-Net generator includes 5 layers of encoders and 5 layers of decoders. Each layer contains a 3×3 convolution, BatchNorm, and the LeakyReLU activation function. The number of channels in each layer of the encoder is 3, 16, 32, 64, 128, and the number of channels in each layer of the decoder is 128, 64, 32, 16, 3. Skip connections are used between the corresponding layers of the encoder and the decoder, and a 1×1 convolution and the sigmoid activation function are used in the last layer of the decoder to output an RGB image.
[0031] The method for generating high-speed moving fabric defect images based on a style transfer network of the present invention is also as follows: In step 3), the training process of the MVU-Style-Transfer network model is specifically as follows:
[0032] 3-1), Input each image set into the MVU-Style-Transfer network model in turn. For each image set, the model generates N images shifted down by 1 pixel based on the static fabric defect images in the image set as the image sequence for synthesizing the intermediate image;
[0033] 3-2), Randomly initialize a weight sequence of length N, send the weight sequence into a randomly initialized multi-layer perceptron MLP. After the weight sequence output by the multi-layer perceptron MLP is normalized by SoftMax, it is weighted and fused with the image sequence to obtain an initial intermediate image. Calculate the feature reconstruction loss between the initial intermediate image and the high-speed moving fabric defect images in the image set, optimize the gradient, and backpropagate to obtain the trained multi-layer perceptron MLP. After another forward process and SoftMax normalization, obtain the final weight sequence. Weight and fuse the weight sequence with the image sequence to obtain the final intermediate image;
[0034] 3-3), Feed the intermediate image into the U-Net generator after random initialization. The output image obtained is the generated image. During the process of training the U-Net generator, for each image set, the static fabric defect images, high-speed moving fabric defect images, and intermediate images within the image set are all randomly cropped using a square with the shorter side length of the image as the side length. When performing the random cropping operation on the static fabric defect images, save the cropping frame parameters i, j, h, w, that is, the starting row coordinate of the cropping area, the starting column coordinate of the cropping area, the height of the cropping area, and the width of the cropping area, and apply them to the cropping process of the intermediate image, so that the cropping position of the intermediate image is the same as that of the static fabric defect image;
[0035] 3-4), Perform standardization processing and Resize on the cropped static fabric defect images, high-speed moving fabric defect images, and intermediate images. The size of the image after Resize is 224×224. Feed the cropped and Resized intermediate image into the U-Net generator to obtain an output image of size 224×224; Then the model uses the self-similarity metric of the deepest layer key of the DINO-ViT pre-trained network to establish a structural loss function based on the cropped and Resized static fabric defect images and the 224×224-sized output image;
[0036] 3-5), Use the L2 norm distance of the CLS token in the deepest layer of the DINO-ViT pre-trained network to establish a visual appearance loss function based on the cropped and Resized high-speed moving fabric defect images and the 224×224-sized output image. Calculate the total loss of the U-Net generator in the MVU-Style-Transfer network model through the structural loss function and the visual appearance loss function, optimize the gradient, and backpropagate until all image sets are traversed and trained, which is recorded as one epoch. Repeat several epochs until the total loss tends to a smaller stable value to obtain the trained MVU-Style-Transfer network model.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. The method for generating high-speed moving fabric defect images based on the style transfer network of the present invention utilizes a multi-layer perceptron MLP to synthesize a static fabric defect image into an intermediate image, and uses a feature reconstruction loss to establish the loss calculation between the intermediate image and the target high-speed fabric defect image for training the multi-layer perceptron MLP, reducing the perceptual gap in the texture motion appearance between the intermediate image and the target image. This method can effectively simulate the texture appearance imaging of the fabric under high-speed movement.
[0039] 2. The high-speed moving fabric defect image generation method based on the style transfer network of the present invention uses U-Net as the generator, adopts a multi-source single-target style transfer framework, trains the generator with a combination of multiple different source images and a single target image to enhance the generalization performance of the generator, obtains weights applicable to multiple source images, ensures the image quality generated based on non-specific source images, and realizes efficient image generation.
[0040] 3. The high-speed moving fabric defect image generation method based on the style transfer network of the present invention utilizes the Dino-ViT pre-trained feature extraction network, uses the self-similarity measure based on the deepest layer key of the network and the L2 norm distance of the CLS token as the structure loss function and the visual appearance loss function, uses the structure loss between the generated image and the static fabric defect image, uses the visual appearance loss between the generated image and the high-speed moving fabric defect image to train the generator, ensures that the structure and visual appearance of the generated image are highly consistent with the static fabric defect image and the high-speed moving fabric defect image respectively, and realizes high-quality image generation.
[0041] 4. The high-speed moving fabric defect image generation method based on the style transfer network of the present invention efficiently generates diverse images with the characteristics of high-speed moving fabric defect images, provides sufficient training samples for the target detection model, thereby improving the detection rate of the model under high-speed operation, and can effectively solve the problem that the detection rate of the target detection model significantly decreases under high-speed motion due to insufficient generalization.
[0042] 5. The high-speed moving fabric defect image generation method based on the style transfer network of the present invention realizes the high-fidelity generation of high-speed moving fabric defect images on the basis of reducing the hardware resources required for training and reducing the dependence on the number of high-speed moving fabric defect images, improves the generalization ability of the target detection model in a high-speed operation environment, injects new impetus into the industrial development, and promotes the intelligent upgrading of traditional industries.
Description of the Drawings
[0043] Figure 1 is the flowchart of the high-speed moving fabric defect image generation method based on the style transfer network of the present invention.
[0044] Figure 2 is the schematic diagram of the texture motion appearance of the fabric defect image in step 2) of the present invention.
[0045] Figure 3 is the schematic diagram of the visual appearance of the fabric defect image in step 2) of the present invention.
[0046] Figure 4 is the schematic diagram of the structure of the fabric defect image in step 2) of the present invention.
[0047] Figure 5 It is a schematic structural diagram of the MVU-Style-Transfer network model in step 3) of the present invention.
[0048] Figure 6 It is the overall effect diagram after generating static fabric defects by using the generation method of the present invention.
Specific Embodiment
[0049] Please refer to the attached drawings of the specification Figure 1 As shown, the present invention is a method for generating high-speed moving fabric defect images based on a style transfer network, which includes the following process steps:
[0050] 1), Use an industrial camera to collect static fabric defect images and high-speed moving fabric defect images, take the static fabric defect images as source images, and the high-speed moving fabric defect images as target images to make a training set.
[0051] In this embodiment, the specific process of this step is:
[0052] 1-1), Use an industrial camera to collect several static fabric images with defects, adjust the speed of the fabric inspection machine to 60m / min, and collect several high-speed moving fabric images with defects.
[0053] 1-2), Perform data augmentation processing on the static fabric images to obtain enhanced static fabric images. After data augmentation of the static fabric defect images, they are combined with the high-speed moving fabric defect images to form a training set. The data augmentation methods include flipping, rotating, adding noise, changing contrast, changing brightness, etc.
[0054] All the collected static fabric images and enhanced static fabric images are extended into a static fabric image dataset, which is divided by defect categories and used as the source image set. Take a single image from the source image set, and use a single high-speed moving fabric image of the same defect category as the target image to construct a single-source single-target image set. Finally, construct a data combination of the same category of defects. The number of image sets in the data combination of each category of defects is not less than 1000, and finally a training set for each category is obtained.
[0055] 2) Feature analysis and decoupling of static and high-speed motion fabric defect images are performed to construct an MVU-Style-Transfer network model. The construction process is as follows: multiple static fabric defect images are used as input images, and high-speed motion fabric defect images are used as target images. Multi-layer perceptron MLP is used in combination with feature reconstruction loss to simulate the texture appearance imaging of fabric under high-speed motion; U-Net is used as the generator, and a multi-source single-target style transfer framework is adopted. In combination with the Dino-ViT pre-trained feature extraction network, the self-similarity metric based on the deepest bond of the network and the L2 norm distance of CLStoken are used as the structural loss function and the visual appearance loss function to train the generator to obtain the MVU-Style-Transfer network.
[0056] Specifically, this step decouples the image into three parts: texture motion appearance, visual appearance, and structure, and builds the MVU-Style-Transfer network model based on these three parts. Texture motion appearance refers to the image of the surface texture of the cloth in the camera when it moves at a certain speed, such as Figure 2 As shown in the figure, data 1 and data 2 correspond to two types of weft fabric defects respectively. Both sets of data include a high-speed motion fabric defect image and a static fabric defect image, as well as a local magnified image of the corresponding defective part. Specifically, when the cloth is in high-speed motion, the texture motion appearance will be streamlined, with the defective part producing a trailing image, while the static fabric defect image will show a clearer yarn arrangement. Visual appearance refers to the difference in imaging features caused by the dynamic reflection and scattering effects generated by the movement of the cloth under the action of a complex ambient light field, such as Figure 3 As shown in the figure, data 1 and data 2 correspond to two types of weft fabric defects respectively. Both sets of data contain a high-speed motion fabric defect image and a static fabric defect image. Specifically, the static defect image and the high-speed motion defect image have significant differences in color or brightness. Structure refers to the relative position of the fabric background and the defect and the geometric shape of the defect. Fabric defects will present different structures due to the quality fluctuations of fabric raw materials and the uncertainty of the production process, such as Figure 4 As shown in Figure 1, Data 1 and Data 2 correspond to two types of weft fabric defects, respectively. Both sets of data contain cloth defect images with different structures. Based on the above three parts of decoupling, the structure of the MVU-Style-Transfer network model is constructed as follows: Figure 5 shown.
[0057] The U-Net generator includes 5 layers of encoders and 5 layers of decoders. Each layer contains a 3×3 convolution, BatchNorm, and a LeakyReLU activation function. The number of channels in each layer of the encoder is 3, 16, 32, 64, 128, and the number of channels in each layer of the decoder is 128, 64, 32, 16, 3. Skip connections are used between the corresponding layers of the encoder and decoder, and a 1×1 convolution and a sigmoid activation function are used in the last layer of the decoder to output an RGB image.
[0058] Using U-Net as the generator, an efficient generation is achieved with a multi-source single-target style transfer framework. U-Net can effectively process the details and multi-scale features of images. The encoder-decoder architecture retains high-resolution detail information through skip connections, enabling it to extract deep features and retain details in the original image simultaneously during the generation process. By adding a 1×1 convolution in the output layer and using a sigmoid activation function to output an RGB image. Common style transfer networks mostly adopt single-source single-target and single-source multi-target transfer methods. After training, the obtained training weights often only target specific image pairs or specific images, resulting in limited generalization performance. By training the generator with combinations of multiple different source images and a single target image, the generalization performance of the generator is enhanced to obtain weights applicable to multiple source images.
[0059] The specific method for texture appearance imaging is as follows: Based on the input static fabric defect image, an image sequence that is shifted down by 1 pixel value in sequence is synthesized. The length of the sequence is equal to the input dimension of the multi-layer perceptron MLP. A weight sequence of the same length is randomly initialized, and the weight sequence is fed into the multi-layer perceptron MLP. The weight sequence output by the multi-layer perceptron MLP is normalized and then weighted and fused with the image sequence to obtain a synthesized intermediate image.
[0060] Furthermore, the multi-layer perceptron MLP includes an input layer, a hidden layer, and an output layer. The dimension of the input layer is calculated from the fabric movement speed and the camera exposure time. The hidden layer consists of three fully connected layers with 128 nodes each and a Relu activation function. The dimension of the output layer is the same as that of the input layer, and SoftMax is used to normalize the output.
[0061] The input dimension of the MLP is calculated from the fabric movement speed and the camera exposure time:
[0062] N = TVk
[0063]
[0064] where N represents the dimension of the input layer and the output layer, T represents the exposure time used by the camera, V represents the fabric movement speed, and k represents the ratio of the camera frame size to the fabric image size; I compositeThe intermediate image representing the MLP output, α n , I n respectively represent the weight sequence of the input layer of the multi-layer perceptron MLP and the image sequence synthesized based on the input static fabric defect image. MLP(·) represents the weight sequence returned after fitting by the multi-layer perceptron, and softmax(·) represents the SoftMax normalization of the weight sequence output by the multi-layer perceptron MLP.
[0065] In this embodiment, the camera frame size is 350mm×235mm, and the size of the captured cloth image is 1920×1080 pixels. Therefore,
[0066]
[0067] The feature reconstruction loss is used to train the multi-layer perceptron MLP. Through the pre-trained VGG network, it compares the feature representations of the intermediate image and the synthesized image of the target high-speed moving fabric defect image at different levels, calculates the L2 norm distance between the two feature maps, and enables the multi-layer perceptron MLP to fit and adjust the weight sequence based on the image sequence, reducing the visual perception gap of the texture motion appearance between the intermediate image and the target image. The texture motion appearance refers to the imaging of the surface texture of the fabric in the camera when the fabric moves at a certain speed.
[0068] The specific calculation formula of the feature reconstruction loss is:
[0069]
[0070] where represents the feature reconstruction loss, specifically referring to the difference in the j-th layer feature space of I composite , I a in the pre-trained network φ, I composite , I a respectively represent the intermediate image output by the MLP and the high-speed moving fabric defect image, φ j (·) represents the feature map of the j-th layer of the VGG network, ||·||2 represents the L2 norm distance between the feature maps, C j , H j , W j respectively represent the number of channels, height, and width of the j-th layer feature map.
[0071] Furthermore, common style transfer models often use convolutional neural networks to extract features through local convolutional kernels. When facing large aspect ratio defects, the feature extraction ability is relatively limited. However, ViT uses a global self-attention mechanism, which can integrate information across regions and capture long-range dependencies, and performs more excellently in extracting large aspect ratio defect features.
[0072] In ViT, an image is segmented into a series of non - overlapping image patches, which are linearly embedded to generate feature vector patch tokens, and positional embeddings are added. The model also introduces a [CLS] token as the global representation of the image. The token sequence containing the image patch features and the [CLS] token is input into a multi - layer Transformer. Each layer consists of a normalization layer LN, a multi - head self - attention MSA module, and an MLP block. Queries, Keys, and Values are calculated through the self - attention mechanism to fuse the information of each image patch. After being processed by all layers, the [CLS] token passes through an additional MLP layer to generate the final image representation.
[0073]
[0074] Among them, represents the output tokens after being processed by the multi - head self - attention (MSA) module in the l - th layer of ViT, T l-1 represents the output tokens of the (l - 1) - th layer of ViT, T l is the output tokens of the image I after being processed by the multi - layer perceptron (MLP) in the l - th layer of ViT; Q l , K l , V l respectively represent the query vector, key vector, and value vector of the l - th layer of ViT, respectively represent the query weight matrix, key weight matrix, and value weight matrix of the l - th layer of ViT.
[0075] The multi - source single - target style transfer framework performs multiple and repeated single - source single - target migrations during training, and trains the generator with combinations of multiple different source images and a single target image, so that the generator has a certain generalization ability.
[0076] Furthermore, DINO - ViT is a ViT model without labels, trained using the self - distillation method. Its deep - layer features capture rich semantic information in the fine - grained space and can effectively distinguish the background and the foreground. Using the self - similarity measure of the deepest - layer keys and the L2 - norm distance of the CLS token in the DINO - ViT pre - trained network as the structural loss function and the visual appearance loss function, where the structural loss function and the visual appearance loss function satisfy the following formula:
[0077]
[0078] Among them, L app represents the visual appearance loss function, represents the [CLS] token extracted in the L - th layer of the ViT model, I a and I orespectively represent the high-speed moving fabric defect image and the image generated by the generator, ||·||2 represents the L2 norm calculation; S L (I) ij represents the self-similarity between patchi and patchj of image I at the L-th layer of the ViT model, and cos-sim(·) represents the cosine similarity calculation. and respectively represent the keys of patchi and patchj at the L-th layer, L structure represents the structure loss function, ||·|| F represents the F norm calculation, I s and I o respectively represent the static fabric defect image and the image generated by the generator.
[0079] 3), use the training set to train the MVU-Style-Transfer network model to obtain the trained MVU-Style-Transfer network model.
[0080] Specifically, the training process of the MVU-Style-Transfer network model is as follows:
[0081] 3-1), input each image set into the MVU-Style-Transfer network model in turn. For each image set, the model generates N images shifted down by 1 pixel based on the static fabric defect image in the image set as the image sequence for synthesizing the intermediate image. In this embodiment, N = 11 is obtained from the above calculation.
[0082] 3-2), randomly initialize a weight sequence of length N. Send the weight sequence into the randomly initialized multi-layer perceptron MLP. After the weight sequence output by the multi-layer perceptron MLP is normalized by SoftMax, it is weighted and fused with the image sequence to obtain an initial intermediate image. Calculate the feature reconstruction loss between the initial intermediate image and the high-speed moving fabric defect image in the image set, optimize the gradient, and backpropagate to obtain the trained multi-layer perceptron MLP. After another forward process and SoftMax normalization, the final weight sequence is obtained. The weight sequence is weighted and fused with the image sequence to obtain the final intermediate image.
[0083] 3-3), Feed the intermediate image into the U-Net generator after random initialization. The output image obtained is the generated image. During the process of training the U-Net generator, for each image set, the static fabric defect images, high-speed moving fabric defect images, and intermediate images within the image set are all randomly cropped using a square with the length of the shorter side of the image as the side length. When performing the random cropping operation on the static fabric defect images, save the cropping frame parameters i, j, h, w, that is, the starting row coordinate of the cropping area, the starting column coordinate of the cropping area, the height of the cropping area, and the width of the cropping area, and apply them to the cropping process of the intermediate image, so that the cropping positions of the intermediate image and the static fabric defect images are the same.
[0084] 3-4), Perform normalization processing and Resize on the cropped static fabric defect images, high-speed moving fabric defect images, and intermediate images. The size of the image after Resize is 224×224. Feed the cropped and Resized intermediate image into the U-Net generator to obtain an output image of size 224×224; then the model uses the self-similarity metric of the deepest layer keys of the DINO-ViT pre-trained network to establish a structure loss function based on the cropped and Resized static fabric defect images and the 224×224-sized output image.
[0085] 3-5), Use the L2 norm distance of the CLS token in the deepest layer of the DINO-ViT pre-trained network to establish a visual appearance loss function based on the cropped and Resized high-speed moving fabric defect images and the 224×224-sized output image. Calculate the total loss of the U-Net generator in the MVU-Style-Transfer network model through the structure loss function and the visual appearance loss function, optimize the gradient, and backpropagate until all image sets are traversed and trained, which is recorded as one epoch. Repeat several epochs until the total loss tends to a smaller stable value to obtain the trained MVU-Style-Transfer network model.
[0086] 4), Input the static fabric defect image to be generated into the trained MVU-Style-Transfer network model to obtain the generation result of the high-speed moving fabric defect image.
[0087] Figure 6It is the effect diagram of generating an image by using the model of the present invention. In order to verify the performance of the proposed generation method of the present invention, an industrial digital camera is used to collect high-speed moving fabric defect images as the original data set. Half of the images in the original data set are used as the test set to evaluate the influence on the training effect of the YOLOv5 model before and after data augmentation. Two fabric defect image data sets are made based on the remaining original data set and the generated images for the training of the YOLOv5 model. The first data set only contains the original data set, and the second data set combines the original data set and the generated images in a quantity ratio of 1:1 to complete data augmentation. Both data sets are surface defect image data sets containing 2 types of weft defects, and the division ratio of the training set and the validation set is 8:2. The comparison of the detection data before and after data augmentation is shown in Table 1.
[0088] In this embodiment, the prediction results are calculated as follows: mAP@.5 means that when the IoU is set to 0.5, the AP of all pictures of each class is calculated, and then the average is taken for all classes; mAP@.5:.95 means the precision averaged for all classes at different IoU thresholds from 0.5 to 0.95 with a step size of 0.05; the precision Precision represents the proportion of the number of accurately found in the prediction; the recall Recall represents the proportion of the correctly predicted in the prediction.
[0089] Table 1 Detection data before and after data augmentation
[0090]
[0091] As can be seen from Table 1, the generation method of the present invention can generate high-quality high-speed moving fabric defect images, and using the images generated by the model for training can significantly improve the detection performance of the detection model.
[0092] In summary, the high-speed moving fabric defect image generation method based on the style transfer network of the present invention preprocesses the data set, including data set image enhancement and data set production; analyzes and decouples the features of the fabric defect images, and constructs a network model based on the decoupled content; in the model, through a multi-layer perceptron MLP combined with a feature reconstruction loss, an intermediate image with motion blur characteristics is synthesized based on static fabric defect images; uses U-Net as the generator, takes the intermediate image as the input, adopts a multi-source single-target style transfer framework, and trains the generator with a combination of various different static fabric defect images and a single high-speed moving fabric defect image to improve the generalization of the generator; combines the Dino-ViT pre-trained feature extraction network, and uses the self-similarity metric of the deepest layer key of the network and the L2 norm distance of the CLS token as the structural loss function and the visual appearance loss function to train the generator, obtaining the MVU-Style-Transfer network. Within different image combinations, it ensures that the structure of the corresponding generated image matches the static fabric defect images within the combination, and at the same time matches the visual appearance of the target high-speed moving fabric defect image, realizing the efficient generation of a brand-new defect image with diverse structures and the characteristics of high-speed moving fabric defect images, providing sufficient training samples for the target detection model, thereby improving the detection rate of the model under high-speed operation, effectively solving the problem of a significant decrease in the detection rate under high-speed motion due to insufficient model generalization, and meeting the actual industrial production requirements.
[0093] The above specific implementation manners are only preferred embodiments of the present creation, and are not intended to limit the present creation. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present creation shall be included within the protection scope of the present creation.
Claims
1. A method for generating high-speed moving fabric defect images based on a style transfer network, characterized in that: The process steps include: 1) Using an industrial camera to collect static fabric defect images and high-speed motion fabric defect images, using static fabric defect images as source images and high-speed motion fabric defect images as target images to create a training set; 2) Feature analysis and decoupling of static and high-speed motion fabric defect images are performed to construct an MVU-Style-Transfer network model. The construction process is as follows: using a multi-layer perceptron MLP and combining feature reconstruction loss to simulate the texture appearance imaging of fabrics under high-speed motion; using U-Net as a generator, adopting a multi-source single-target style transfer framework, combined with the Dino-ViT pre-trained feature extraction network, using the self-similarity metric based on the deepest bond of the network and the L2 norm distance of CLStoken as the structural loss function and visual appearance loss function to train the generator, and obtain the MVU-Style-Transfer network; 3) Using the training set to train the MVU-Style-Transfer network model, and obtaining a trained MVU-Style-Transfer network model; 4) Input the static fabric defect image to be generated into the trained MVU-Style-Transfer network model to obtain the generation result of the high-speed motion fabric defect image.
2. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 1, characterized in that: The step 1) is specifically as follows: 1-1), use an industrial camera to capture several static fabric images with defects, adjust the speed of the fabric inspection machine to 60m / min, and capture several high-speed moving fabric images with defects; 1-2), perform data enhancement processing on the static fabric image to obtain an enhanced static fabric image; expand all the collected static fabric images and enhanced static fabric images into a static fabric image dataset, divide them according to the defect category and use them as a source image set, take a single image in the source image set, and use a single high-speed motion fabric image of the same defect category as the target image to construct a single-source and single-target image set, and finally construct a data combination of defects of the same category. The number of image sets in the data combination of each defect category is not less than 1000, and finally a training set for each category is obtained.
3. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 2, characterized in that: In step 1-2), the data enhancement method includes flipping, rotating, adding noise, changing contrast, and changing brightness.
4. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 1, characterized in that: In the step 2), the image is specifically decoupled into three parts: texture motion appearance, visual appearance and structure, and the MVU-Style-Transfer network model is constructed based on these three parts.
5. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 1, characterized in that: In the step 2), the specific method of texture appearance imaging is: synthesizing an image sequence that is shifted down by 1 pixel value in sequence based on the input static fabric defect image, wherein the sequence length is equal to the input dimension of the multi-layer perceptron MLP, randomly initializing a set of weight sequences of the same length, sending the weight sequence to the multi-layer perceptron MLP, and normalizing the weight sequence output by the multi-layer perceptron MLP and weighted fusion with the image sequence to obtain a synthesized intermediate image; The input dimension of MLP is calculated by the fabric movement speed and camera exposure time: N=TVk Where N represents the dimensions of the input layer and the output layer, T represents the exposure time used by the camera, V represents the speed of the fabric movement, and k represents the ratio of the camera frame to the fabric image size; I composite represents the intermediate image output by MLP, α n , I n They represent the weight sequence of the input layer of the multi-layer perceptron MLP and the image sequence synthesized based on the input static fabric defect image, respectively. MLP(·) represents the weight sequence returned after the multi-layer perceptron fitting. Softmax(·) represents the SoftMax normalization of the weight sequence output by the multi-layer perceptron MLP.
6. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 1, characterized in that: In the step 2), the calculation formula of the feature reconstruction loss is: in Represents feature reconstruction loss, specifically I composite , I a The difference in the j-th layer feature space in the pre-trained network φ, I composite , I a They represent the intermediate image output by MLP and the high-speed motion fabric defect image, respectively, j (·) represents the feature map of the jth layer of the VGG network, ||·||2 represents the L2 norm distance between feature maps, and C j , H j , W j They represent the number of channels, height, and width of the j-th layer feature map, respectively.
7. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 1, characterized in that: In step 2), the structural loss function and the visual appearance loss function satisfy the following formula: Among them, L app represents the visual appearance loss function, represents the [CLS] token extracted at the Lth layer of the ViT model, I a and I o represent the high-speed motion fabric defect image and the image generated by the generator, respectively, ||·||2 represents the L2 norm calculation; S L (I) ij represents the self-similarity between patchchi and patchj of image I in the Lth layer of the ViT model, cos-sim(·) represents the cosine similarity calculation, and Represent the keys of patchi and patchj at layer L, respectively. structure represents the structural loss function, ||·|| F Indicates the F-norm calculation, I s and I o Represent the static fabric defect images and the images generated by the generator, respectively.
8. The method for generating high-speed moving fabric defect images based on a style transfer network as claimed in claim 1, characterized in that: In the step 2), the multi-layer perceptron MLP includes an input layer, a hidden layer and an output layer. The dimension of the input layer is calculated by the fabric movement speed and the camera exposure time. The hidden layer consists of three nodes with 128 nodes and a Relu fully connected layer as the activation function. The dimension of the output layer is consistent with the input layer, and SoftMax is used to normalize the output.
9. The method for generating high-speed moving fabric defect images based on a style transfer network as claimed in claim 1, characterized in that: The U-Net generator includes 5 layers of encoders and 5 layers of decoders, each layer contains 3×3 convolution, BatchNorm and LeakyReLU activation functions, the number of channels of each layer of the encoder is 3, 16, 32, 64, 128, and the number of channels of each layer of the decoder is 128, 64, 32, 16, 3. Jump connections are used between the corresponding layers of the encoder and decoder. The last layer of the decoder uses 1×1 convolution and sigmoid activation function to output RGB images.
10. The method for generating high-speed moving fabric defect images based on a style transfer network according to claim 5, characterized in that: In step 3), the training process of the MVU-Style-Transfer network model is specifically as follows: 3-1), each image set is sequentially input into the MVU-Style-Transfer network model. For each image set, the model generates N images shifted down by 1 pixel based on the static fabric defect images in the image set as an image sequence for synthesizing the intermediate image; 3-2), randomly initialize a set of weight sequences of length N, send the weight sequence to the randomly initialized multi-layer perceptron MLP, the weight sequence output by the multi-layer perceptron MLP is normalized by SoftMax, and then weighted fused with the image sequence to obtain an initial intermediate image, calculate the feature reconstruction loss of the initial intermediate image and the high-speed motion fabric defect image in the image set, optimize the gradient, and back propagate to obtain the trained multi-layer perceptron MLP, and then go through a forward process and SoftMax normalization to obtain the final weight sequence, and weighted fused the weight sequence with the image sequence to obtain the final intermediate image; 3-3), the intermediate image is sent to the randomly initialized U-Net generator, and the output image obtained is the generated image. In the process of training the U-Net generator, for each image set, the static fabric defect image, the high-speed motion fabric defect image and the intermediate image in the image set are randomly cropped with a square with the short side of the image as the side length. When the random cropping operation is performed on the static fabric defect image, the cropping frame parameters i, j, h, w, i.e., the starting row coordinates of the cropping area, the starting column coordinates of the cropping area, the height of the cropping area, and the width of the cropping area are saved, and applied to the cropping process of the intermediate image, so that the cropping position of the intermediate image is consistent with that of the static fabric defect image; 3-4), the cropped static fabric defect image, high-speed motion fabric defect image and intermediate image are standardized and resized. The resized image size is 224×224. The cropped and resized intermediate image is sent to the U-Net generator to obtain an output image of 224×224 size; then the model uses the self-similarity metric of the deepest bond of the DINO-ViT pre-trained network to establish a structural loss function based on the cropped and resized static fabric defect image and the output image of 224×224 size; 3-5), using the L2 norm distance of the deepest CLStoken of the DINO-ViT pre-trained network, a visual appearance loss function is established based on the cropped and resized high-speed motion fabric defect image and the output image of 224×224 size, and the total loss of the U-Net generator in the MVU-Style-Transfer network model is calculated through the structural loss function and the visual appearance loss function. The gradient is optimized and back-propagated until all image sets are traversed and trained, which is recorded as an epoch. Several epochs are repeated until the total loss tends to a smaller stable value, and the trained MVU-Style-Transfer network model is obtained.
Citation Information
Cited By
Fuzzing and pilling rating method and device based on visual continuous regression
CN121810682A