Training Method and Related Device for Silk Thread Fault Detection Model in Spinning Process

By performing discrete cosine transformation and multi-layer neural network processing on the spinning box image, multi-layer features are generated to detect wire failures during spinning, solving the problem of low detection efficiency in the prior art, achieving efficient and accurate wire failure detection, and improving production quality and efficiency.

CN120279346BActive Publication Date: 2025-08-05ZHEJIANG HENGYI PETROCHEMICAL CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510783720.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-05
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In spinning process, it is difficult for the prior art to efficiently and accurately detect wire failures, such as broken wires and floating wires, resulting in low production efficiency and product quality problems.

Method used

Using artificial intelligence technology, the images collected by spinning box are discrete cosine transformed, frequency domain representation is extracted, and neural network modules such as zero-frequency enhancers, low-frequency reconstructors and high-frequency refiners are used to generate multi-level features, and finally fault judgment is performed through classifiers to optimize model parameters to improve detection accuracy.

Benefits of technology

It realizes efficient and accurate detection of wire failures, improves the quality control and production efficiency of spinning production, and reduces production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279346B_ABST
    Figure CN120279346B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method and related device for a yarn fault detection model in a spinning process. It relates to the field of artificial intelligence technology. The method comprises: performing discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; inputting the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a model to be trained to obtain a first global feature of the sample image; inputting the frequency domain representation into a low-frequency regenerator of a model to be trained to obtain an intermediate semantic feature of the sample image; inputting the frequency domain representation and the first global feature into a high-frequency refiner of a model to be trained to obtain a fine-grained feature of the sample image; inputting the first global feature, the intermediate semantic feature and the fine-grained feature into a fusion device of a model to be trained to obtain a fused feature; inputting the fused feature into a classifier of a model to be trained to obtain a classification result for the sample image; optimizing the model parameters of the model to be trained based on the classification result to obtain a yarn fault detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as spinning technology and deep learning. Background Art

[0002] In the industrial spinning process, yarn quality directly determines the performance and quality of the final product. As the core output of the spinning process, the quality of yarn plays a decisive role in key performance indicators of downstream textiles, such as strength, toughness, and dyeing consistency. Summary of the Invention

[0003] The present disclosure provides a method for training a yarn fault detection model in a spinning process to solve or alleviate one or more technical problems in related technologies.

[0004] In a first aspect, the present disclosure provides a method for training a yarn fault detection model in a spinning process, comprising:

[0005] Performing discrete cosine transform on the sample image to obtain a frequency domain representation of the sample image; the sample image is obtained based on image acquisition of the spinning beam;

[0006] Inputting the zero-frequency component in the frequency domain representation and the sample image into the zero-frequency enhancer of the model to be trained to obtain the first global feature of the sample image;

[0007] Input the frequency domain representation into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image;

[0008] Inputting the frequency domain representation and the first global feature into the high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image;

[0009] Input the first global feature, the intermediate semantic feature, and the fine-grained feature into the fusion device of the model to be trained to obtain the fusion feature;

[0010] The fused features are input into the classifier of the model to be trained to obtain the classification results for the sample image, including the results of floating silk discrimination and broken silk discrimination;

[0011] Based on the classification results, the model parameters of the model to be trained are optimized to obtain a wire fault detection model.

[0012] In a second aspect, the present disclosure provides a training device for a yarn fault detection model in a spinning process, comprising:

[0013] A frequency domain decomposition module is used to perform discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; the sample image is obtained based on image acquisition of a spinning beam;

[0014] a zero-frequency enhancement module, configured to input the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of the model to be trained, thereby obtaining a first global feature of the sample image;

[0015] The low-frequency reconstruction module is used to input the frequency domain representation into the low-frequency reconstruction of the model to be trained to obtain the intermediate semantic features of the sample image;

[0016] A high-frequency refinement module, configured to input the frequency domain representation and the first global feature into a high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image;

[0017] A fusion module is used to input the first global feature, the intermediate semantic feature and the fine-grained feature into the fuser of the model to be trained to obtain a fused feature;

[0018] The classification module is used to input the fusion features into the classifier of the model to be trained to obtain the classification results for the sample image, including the results of the floating silk discrimination and the broken silk discrimination;

[0019] The optimization module is used to optimize the model parameters of the to-be-trained model based on the classification results to obtain a wire fault detection model.

[0020] In a third aspect, the present disclosure provides a method for detecting yarn faults in a spinning process, comprising:

[0021] Capturing images of the spinning beam to obtain images to be processed;

[0022] The image to be processed is input into a yarn fault detection model obtained by a training method based on a yarn fault detection model in a spinning process to detect whether the yarn has floating or broken yarn problems.

[0023] In a fourth aspect, the present disclosure provides a device for detecting yarn faults in a spinning process, comprising:

[0024] An acquisition module is used to acquire images of the spinning box to obtain images to be processed;

[0025] The detection module is used to input the image to be processed into a silk thread fault detection model obtained by a training method based on a silk thread fault detection model in a spinning process, so as to detect whether the silk thread has floating or broken silk threads.

[0026] According to a fifth aspect, an electronic device is provided, including:

[0027] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0028] In a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0029] In a seventh aspect, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.

[0030] In the disclosed embodiments, the advantages of each feature are leveraged to enable the trained model to obtain a more comprehensive and rich image feature representation, thereby more accurately describing the sample image and providing stronger feature support for subsequent classification. Based on the classification results, the parameters of the trained model are continuously optimized, allowing it to gradually learn how to better extract features and perform classification. This results in a wire fault detection model with high classification accuracy on the training dataset.

[0031] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0033] Figure 1 is a flow chart of a method for training a yarn fault detection model in a spinning process according to an embodiment of the present disclosure;

[0034] Figure 2 is a schematic diagram of a process for obtaining a first global feature of a sample image according to an embodiment of the present disclosure;

[0035] Figure 3 is a schematic diagram of a process for obtaining intermediate semantic features of a sample image according to an embodiment of the present disclosure;

[0036] Figure 4 is an exemplary mask feature map according to an embodiment of the present disclosure;

[0037] Figure 5 is a schematic diagram of a process for obtaining fine-grained features of a sample image according to an embodiment of the present disclosure;

[0038] Figure 6 is a schematic diagram of a process for obtaining fusion features according to an embodiment of the present disclosure;

[0039] Figure 7 is a schematic diagram of a process for optimizing model parameters of a model to be trained according to an embodiment of the present disclosure;

[0040] Figure 8 1 is an overall structural diagram of a method for training a yarn fault detection model in a spinning process according to an embodiment of the present disclosure;

[0041] Figure 9 is a schematic flow chart of a method for detecting yarn faults in a spinning process according to one embodiment of the present disclosure;

[0042] Figure 10 2 is a schematic structural diagram of a training device for a yarn fault detection model in a spinning process according to an embodiment of the present disclosure;

[0043] Figure 11 is a schematic structural diagram of a yarn fault detection device in a spinning process according to an embodiment of the present disclosure;

[0044] Figure 12 It is a block diagram of an electronic device used to implement the training method of the yarn fault detection model in the spinning process and / or the yarn fault detection method in the spinning process according to the embodiments of the present disclosure. DETAILED DESCRIPTION

[0045] The following is a further detailed description of the method for training a yarn fault detection model in a spinning process provided by the present disclosure, with reference to the accompanying drawings. Identical reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are illustrated in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0046] In addition, in order to better illustrate the technical solutions of the embodiments of the present disclosure, numerous specific details are provided in the following detailed description. Those skilled in the art will understand that the present disclosure can be implemented without certain specific details. In some examples, methods, means, components, circuits, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.

[0047] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0048] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order between multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.

[0049] As the chemical fiber industry's requirements for product quality continue to increase and production scale continues to expand, quality control in the silk thread production process becomes increasingly critical.

[0050] During the actual spinning process, yarn breakage and drifting are unavoidable. Breakage refers to the sudden breakage of a yarn during the spinning process, which can affect spinning continuity and production efficiency. If this leads to production interruptions and the generation of large amounts of waste yarn, this can cause significant losses to the company. Drifting occurs when a yarn floats irregularly during the spinning process. This can lead to product quality issues and direct losses for the company. Once these abnormalities occur, they can have negative consequences.

[0051] Therefore, there is an urgent need for an efficient, accurate and adaptable yarn fault detection technology that can detect abnormal conditions such as broken yarns and floating yarns in a timely and accurate manner to ensure the quality of spinning products, improve production efficiency, reduce production costs, and meet the growing production needs of the modern textile industry.

[0052] In view of this, in the embodiments of the present disclosure, a training method for a yarn fault detection model in a spinning process and a yarn fault detection method in a spinning process are proposed based on artificial intelligence technology.

[0053] It should be noted that the main types of spun yarns involved in the embodiments of the present disclosure may include one or more of partially oriented yarns (POY), fully drawn yarns (FDY), polyester staple fibers, etc. For example, the types of yarns may specifically include polyester partially oriented yarns, polyester fully drawn yarns, polyester drawn yarns, polyester staple fibers, etc.

[0054] The following describes the training method of the thread fault detection model and the thread fault detection method respectively.

[0055] like Figure 1FIG. 1 is a schematic diagram of a processing flow of a method for training a wire fault detection model provided by an embodiment of the present disclosure, which can be specifically implemented as follows:

[0056] S101 , performing discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image.

[0057] Sample images are acquired by capturing images of the spinning beam and represent the state of the yarn at different moments. After acquiring the sample images, to improve model training efficiency and test results, they are first standardized and resized to a uniform size. The processed sample images are then segmented into n×n pixel blocks to provide data for the subsequent discrete cosine transform (DCT), ensuring the standardization and consistency of the processing. Here, n is an integer greater than 1, such as 8. The specific value depends on the actual business needs and the requirements of the DCT transform.

[0058] The discrete cosine transform (DCT) is a transform technique related to the Fourier transform that is used to convert signals from one domain to the frequency domain. In the disclosed embodiment, by performing a discrete cosine transform on a sample image, the sample image can be converted from the spatial domain to frequency domains at different levels, thereby obtaining a frequency domain representation of the sample image. For example, the zero-frequency component in the frequency domain representation can capture global information such as the overall brightness and contrast of the sample image; the low-frequency component in the frequency domain representation can reflect slowly changing areas in the sample image, such as portions of the sample image where the silk threads are more evenly distributed; and the high-frequency component in the frequency domain representation reflects rapidly changing details in the sample image, such as the edge contours of the silk threads and the interlaced texture of the silk threads in the sample image.

[0059] In specific implementation, the frequency domain representation of the sample image can be obtained through formula (1), as follows:

[0060]

[0061] In formula (1), Represents the frequency domain representation of the sample image, that is, c Channel, h Block row, w DCT coefficients of block columns; Represents a sample image x Middle c Channel, nh+i Row, No. nw+j The pixel value of the column; It is the cosine basis function in the two-dimensional DCT transform, which is used to convert the pixel value in the spatial domain into the frequency domain coefficient. The cosine basis function is used to calculate the frequency component in the horizontal direction; Used to calculate the frequency component in the vertical direction, u,v are frequency indices, ranging from 0 to n-1; i,j are the pixel indices within the pixel block, ranging from 0 to n-1.

[0062] S102: Input the zero-frequency component in the frequency domain representation and the sample image into the zero-frequency enhancer of the model to be trained to obtain a first global feature of the sample image.

[0063] In the frequency domain, the zero-frequency component usually refers to the frequency index value u =0 and v =0, the corresponding frequency domain component, thus, the zero-frequency component can be obtained based on formula (2):

[0064]

[0065] In formula (2), Represents the zero-frequency component. Other parameters are the same as those in formula (1) and will not be repeated here.

[0066] The zero-frequency component and the sample image in the frequency domain representation are input into the zero-frequency enhancer of the model to be trained. Based on the overall information of the sample image provided by the zero-frequency component and the information provided by the sample image, the zero-frequency enhancer generates a first global feature containing global information of multiple dimensions of the sample image. In other words, the zero-frequency enhancer is the first neural network module of the model to be trained, which is used to extract the first global feature from the zero-frequency component and the sample image.

[0067] S103, inputting the frequency domain representation into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image.

[0068] Among them, the low-frequency regenerator is the second neural network module of the model to be trained, which is used to extract the features of the low-frequency components from the frequency domain representation to obtain the mid-level semantic features.

[0069] Alternatively, the frequency domain representation is fed into the low-frequency reconstruction of the model to be trained. This reconstruction processes the low-frequency components of the frequency domain representation and extracts mid-level semantic features of the sample image. These mid-level semantic features help the model to be trained understand the continuity characteristics of the silk threads in the sample image, such as the areas where the threads are evenly distributed.

[0070] S104: Input the frequency domain representation and the first global feature into the high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image.

[0071] Among them, the high-frequency refiner is the third neural network module of the model to be trained, which is used to extract the features of the high-frequency components from the frequency domain representation and the first global features to obtain the detailed features of the high-frequency components.

[0072] The frequency domain representation and the first global feature are fed into the high-frequency refiner of the model to be trained. The refiner combines the high-frequency components of the frequency domain representation with the first global feature to further extract detailed information from the sample image. These fine-grained features help the model to be trained capture the high-frequency vibration trajectory characteristics of the thread, such as the jittery edges of the thread and the specific form of motion blur.

[0073] S105: Input the first global feature, the intermediate semantic feature, and the fine-grained feature into the fusion device of the model to be trained to obtain a fusion feature.

[0074] The first global feature, intermediate semantic feature and fine-grained feature are input into the fusion device of the model to be trained. The fusion device can combine the first global feature, intermediate semantic feature and fine-grained feature, make full use of the advantages of features at each level, and generate a fusion feature with multi-semantic level feature representation.

[0075] S106: Input the fused features into the classifier of the model to be trained to obtain a classification result for the sample image, wherein the classification result includes a floating silk discrimination result and a broken silk discrimination result.

[0076] The classifier can determine whether there are floating or broken silk threads in the sample image by analyzing the fusion features, and output the corresponding judgment results based on this, that is, clearly indicating whether the silk thread in the current sample image is floating, broken, or in a normal state.

[0077] S107 , based on the classification result, optimizing the model parameters of the model to be trained to obtain a wire fault detection model.

[0078] During implementation, the classification results obtained based on the above classifier can be compared with the corresponding true labels, and indicators such as loss values can be calculated. Then, the model parameters of the model to be trained can be optimized based on this, so that the model to be trained can continuously improve the accuracy of the reasoning process, and finally an effective silk thread fault detection model is obtained, which can be used for silk thread fault detection on newly collected sample images in the actual spinning process.

[0079] In the embodiment of the present disclosure, a discrete cosine transform is performed on the sample image to obtain a frequency domain representation of the sample image, which can convert the sample image collected based on the spinning box from the spatial domain to the frequency domain, so as to facilitate the subsequent targeted processing of different features corresponding to different frequency components. The zero-frequency component in the frequency domain representation and the sample image are input into the zero-frequency enhancer of the model to be trained to obtain the first global feature of the sample image, which can help the model to be trained to extract the features of the sample image as a whole. The frequency domain representation is input into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image, which can help the model to be trained to understand the continuity features of the silk threads in the sample image, thereby providing a basis for detecting silk thread faults. The frequency domain representation and the first global feature are input into the high-frequency refiner of the model to be trained to obtain the fine-grained features of the sample image, which can help the model to be trained to capture more subtle changes in the silk threads in the sample image. The first global features, mid-level semantic features, and fine-grained features are input into the fusion unit of the model to be trained to obtain fused features. By fusing features of different dimensions and levels, the advantages of each feature can be fully utilized, allowing the model to obtain a more comprehensive and rich image feature representation, thereby more accurately describing the sample image and providing stronger feature support for subsequent classification. Based on the classification results, the parameters of the model to be trained are continuously optimized, allowing it to gradually learn how to better extract features and perform classification, resulting in a high classification accuracy rate for the final wire fault detection model on the training dataset.

[0080] In an embodiment of the present disclosure, the zero-frequency enhancer includes a first global average pooling layer and a convolutional layer.

[0081] The global average pooling layer is a pooling operation commonly used in deep learning. It performs a global average pooling operation on the input feature map, averaging all elements in each feature channel to obtain a value that represents the overall information of the feature channel. In the disclosed embodiment, the first global average pooling layer receives a sample image as input and outputs the globally averaged pooled features of the sample image.

[0082] The convolutional layer is an important structure for extracting features in deep learning. By sliding the convolution kernel over the input data and performing convolution operations, it can automatically learn local features and patterns in the data. In the disclosed embodiment, the convolutional layer receives the spliced feature as input and outputs the first global feature.

[0083] During implementation, the zero-frequency component in the frequency domain representation and the sample image are input into the zero-frequency enhancer of the model to be trained to obtain the first global feature of the sample image, such as Figure 2 As shown, it can be implemented as follows:

[0084] S201: Input the sample image into the first global average pooling layer to obtain the global average pooling features of the sample image.

[0085] The global average pooling feature is to find the average value of each channel of the sample image in the spatial dimension (usually height and width), and its calculation method is shown in formula (3):

[0086]

[0087] In formula (3), Represents the global average pooling features of the sample image; Represents a sample image x In the channel dimension c , spatial dimension i (height direction) and j Elements in the width direction; h and w Represents the size of the tensor in height and width dimensions respectively.

[0088] S202: Concatenate the zero-frequency component and the global average pooling feature to obtain a concatenated feature.

[0089] During implementation, the zero-frequency component in the frequency domain representation and the global average pooling feature obtained in the previous step can be spliced in the channel dimension to obtain the spliced feature.

[0090] S203: Input the concatenated features into the convolution layer to obtain the first global features.

[0091] The spliced features are input into the convolution layer for convolution operation, that is, the final first global feature is generated by fusing and reducing the dimensions of the spliced features between channels.

[0092] During implementation, the process of obtaining the first global feature can be expressed by formula (4):

[0093]

[0094] In formula (4), Represents the global average pooling features of the sample image; Represents the zero-frequency component in the frequency domain representation; Represents the global average pooling features of the sample image; Represents the splicing operation of the channel dimension, which is used to splice the zero-frequency component in the frequency domain representation and the global average pooling feature of the sample image; Represents a 1×1 convolution operation, which is used to extract the first global feature of the sample image without changing the spatial size of the feature map.

[0095] In the embodiment of the present disclosure, through global average pooling, the feature map on each channel of the sample image can be averaged and converted into a scalar value, reducing the dimension of the feature, thereby helping the model to be trained to reduce the computational complexity of the model. The zero-frequency component and the global average pooling feature are spliced to obtain a spliced feature, which can integrate the brightness information and global features of the sample image, so that the model to be trained can simultaneously consider the brightness and overall features of the sample image, thereby improving the model's ability to understand the sample image. The spliced feature is input into the convolutional layer, thereby fusing the global brightness and the first global feature, and converting them to a higher-level feature space for representation, which is conducive to obtaining a first global feature that can better reflect the overall features of the sample image, and can provide a more effective feature representation for subsequent fault classification, thereby improving the accuracy of wire fault detection.

[0096] In the disclosed embodiment, the low frequency regenerator includes a learnable frequency domain selector and a normalization layer.

[0097] The learnable frequency domain selector learns the correlation patterns between different frequency bands in the frequency domain representation, thereby generating a frequency weight matrix for frequency selection. This matrix automatically determines which frequency bands' features are most conducive to extracting appropriate mid-level semantic features. The learnable frequency domain selector is a key component for filtering and integrating frequency domain information. Through training and optimization, it can learn patterns for classifying and identifying thread faults, thereby improving the accuracy of thread fault detection.

[0098] The normalization layer normalizes the intermediate features to make their data distribution more stable, which helps the subsequent feature extraction and model training process. It can eliminate the differences in dimensions and numerical ranges between different features and accelerate the convergence of the model.

[0099] During implementation, the frequency domain representation is input into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image, such as Figure 3 As shown, it can be implemented as follows:

[0100] S301 , learning a correlation pattern between different frequency bands in a frequency domain representation based on a learnable frequency domain selector, and generating a frequency weight matrix for selecting frequencies based on the correlation pattern.

[0101] The frequency domain representation is fed into a learnable frequency selector. This selector uses a learning algorithm to analyze the correlation patterns between different frequency bands in the frequency domain representation. For example, in the frequency domain representation of a silk thread sample image, low-frequency bands may be associated with the thread's continuity characteristics, while high-frequency bands may be associated with local texture. Based on these correlation patterns, the learnable frequency selector calculates the importance weight of each frequency component and generates a frequency weight matrix.

[0102] In implementation, generating a frequency weight matrix for selecting frequencies based on the correlation pattern can be achieved based on the following steps:

[0103] Step A1: input the frequency domain representation into the first activation layer of the learnable frequency domain selector to obtain activation features; the activation features are used to represent the correlation pattern between different frequency bands in the frequency domain representation;

[0104] The frequency domain representation is input into the first activation layer of the learnable frequency domain selector. The first activation layer performs a nonlinear transformation on the frequency domain representation through the activation function to obtain activation features. The activation features can reflect the correlation pattern between different frequency bands in the frequency domain representation.

[0105] In step A2, the activation features are input into the second activation layer of the learnable frequency domain selector to obtain a frequency weight matrix.

[0106] The activation features obtained above are input into the second activation layer of the learnable frequency domain selector. The second activation layer undergoes another nonlinear transformation using the activation function, ultimately generating a frequency weight matrix. Each element in the frequency weight matrix corresponds to a frequency component in the frequency domain representation, and the value in the frequency weight matrix represents the importance of that frequency component in constructing the mid-level semantic features. Frequency components with higher weights play a greater role in generating mid-level semantic features, while frequency components with lower weights may be suppressed or ignored.

[0107] In the embodiment of the present disclosure, the learnable frequency domain selector can be expressed as shown in formula (5):

[0108]

[0109] In formula (5), represents the first activation layer, W1 is the weight of the first activation layer, b1 is the bias of the first activation layer. The first activation layer is represented by the frequency domain of the input Projected into a new feature space, and then learned W1 parameters to capture the correlation patterns between different frequency bands; b1 The bias term is used to adjust the distribution of features to help the model better fit the data; It is a nonlinear activation function that can introduce nonlinear characteristics, enabling the model to learn complex feature relationships. Represents the second activation layer, which is used to learn weights based on the activation features obtained based on the first activation layer W2 The parameters of the second activation layer generate a frequency selection weight matrix, which is used to weight different frequency components; b2 The bias term is used to adjust the distribution of the second layer features.

[0110] In the disclosed embodiment, a learnable frequency domain selector first inputs the frequency domain representation into the first activation layer. Using nonlinear transformations, it extracts the correlation pattern between different frequency bands in the frequency domain representation, generating activation features representing these correlation patterns. The activation features are then input into the second activation layer to generate a frequency weight matrix for frequency selection. This matrix adaptively assigns weights to different frequencies, enabling the trained model to adaptively select the frequencies most helpful for yarn fault detection, adapting to different spinning scenarios and fault types. Compared to fixed frequency selection methods, this adaptive approach is more flexible and effective, and can improve the generalization ability of the trained model.

[0111] During the training phase, the weights and biases of the first and second activation layers in the learnable frequency domain selector are continuously optimized so that the learnable frequency domain selector can learn to capture important low-frequency component information. After the training, the weights and biases of the first and second activation layers no longer change, and the frequency domain representation can be adaptively adapted to screen out appropriate low-frequency components for constructing intermediate semantic features.

[0112] S302: Perform an inverse discrete cosine transform on the frequency weight matrix to obtain a first spatial domain feature.

[0113] That is, the generated frequency weight matrix is subjected to an inverse discrete cosine transform, which is converted from the frequency domain back to the spatial domain to obtain the first spatial domain feature.

[0114] S303: Determine intermediate features based on the spatial domain features and the mask features; the mask features include frequency band masks corresponding to each frequency in the frequency domain representation.

[0115] Based on the first spatial domain feature obtained above and the pre-set mask feature, the intermediate feature is determined. The mask feature includes the frequency band mask corresponding to each frequency in the frequency domain representation, such as Figure 4 The figure shows an exemplary mask matrix (i.e., mask feature) provided for an embodiment of the present disclosure. The upper left corner region serves as a band mask corresponding to low frequencies, capturing macroscopic features such as the silk thread outline in the sample image; the middle region serves as a band mask corresponding to medium frequencies, extracting scale texture information from the silk thread in the sample image; and the lower left corner region serves as a band mask corresponding to high frequencies, preserving finer image details in the sample image. Band masks allow for further refinement in filtering frequency components, clarifying which frequency components should be retained or highlighted and which should be suppressed or discarded. For example, the first k×k low-frequency components in the frequency domain representation are retained, while other high-frequency components are masked.

[0116] S304: Input the intermediate features into the normalization layer to obtain intermediate semantic features.

[0117] The determined intermediate features are input into the normalization layer for normalization. The normalization layer normalizes the intermediate feature data to meet specific distribution requirements, thereby obtaining intermediate semantic features.

[0118] In specific implementation, the process of obtaining intermediate semantic features can be described by formula (6), as follows:

[0119]

[0120] In formula (6), Fmid Represents the mid-level semantic features of the sample image; Mk Represents the front of the frequency domain representation k × k Mask features of low-frequency components; k Is a positive integer used to control the range of retained low-frequency components; represents the obtained frequency domain representation; represents the frequency weight matrix for selecting frequencies based on the correlation pattern; represents the inverse discrete cosine transform, which is used to convert the frequency weight matrix from the frequency domain representation back to the spatial domain features; Represents element-by-element multiplication (Hadamard product), which multiplies the elements of corresponding positions of two matrices; LayerNorm () represents the normalization layer, which is used to process the intermediate features and obtain the intermediate semantic features.

[0121] In the disclosed embodiment, the learnable frequency domain selector can deeply explore the complex correlation patterns between different frequency bands in the frequency domain representation, learn the correlation patterns based on the learnable frequency domain selector and generate a frequency weight matrix, so that the matrix can adaptively assign corresponding weights to each frequency, thereby enabling the model to adaptively select the frequencies that are most helpful for silk thread fault detection. The intermediate features are determined based on the spatial domain features and the mask features, so that the intermediate features not only reflect the important features learned based on the frequency correlation in the spatial domain, but also meet the frequency screening requirements specified by the mask features. Standardizing the intermediate features through the normalization layer helps to improve the training stability and convergence speed of the model to be trained, and avoid the training difficulties caused by excessive differences in feature scales. At the same time, the normalized intermediate semantic features are more conducive to subsequent model processing and analysis, thereby improving the model's ability to detect silk thread faults in sample images.

[0122] In the disclosed embodiment, the high-frequency refiner includes a dynamic convolution layer. Unlike traditional convolution layers, the convolution kernel weights in a dynamic convolution layer are not fixed but are dynamically generated based on input features. This enables the model to adaptively adjust the convolution operation based on the characteristics of different sample images, thereby more accurately extracting fine-grained features.

[0123] During implementation, the frequency domain representation and the first global feature are input into the high-frequency refiner of the model to be trained to obtain the fine-grained features of the sample image, such as Figure 5 As shown, it can be implemented as follows:

[0124] S501, extracting high-frequency components from the frequency domain representation.

[0125] Based on the content described in step S303 above, the mask feature retains the front × The low-frequency components of the image are masked out, while other high-frequency components are masked out. Therefore, the high-frequency components can be extracted from the frequency domain representation by selecting the complement of the mask feature to mask out the low-frequency components and retain only the high-frequency components.

[0126] S502: Convert the high-frequency component to the spatial domain to obtain a second spatial domain feature.

[0127] The high-frequency components to be extracted are converted to the spatial domain through inverse discrete cosine transform to obtain the second spatial domain features.

[0128] S503: Determine a convolutional layer weight of the dynamic convolutional layer based on the first global feature.

[0129] The first global feature contains the overall information of the sample image. By using this global information to dynamically generate convolutional layer weights, the convolution operation can be made more targeted. For example, if the first global feature indicates that a certain area in the sample image may contain thread fault characteristics, the generated convolutional layer weights may focus more on extracting details in this area.

[0130] That is, in order to enable the dynamic convolution layer to adaptively adjust according to the overall information of the sample image, so as to better capture the context and semantic information of the sample image, the convolution layer weight of the dynamic convolution layer can be determined based on the global features. The processing process is shown in formula (7):

[0131]

[0132] In formula (7), W dy Represents the convolution layer weights of the dynamic convolution layer; F global represents the first global feature; W a Represents the weight matrix, which is used to perform dot product operation with the first global feature; Softmax () represents a function that converts the input vector into a probability distribution (such as an activation layer) and is used to generate the convolutional layer weights.

[0133] S504: Input the second spatial domain feature into a dynamic convolution layer using convolution layer weights to obtain fine-grained features.

[0134] That is, the second spatial domain feature is input into the dynamic convolution layer using the convolution layer weights generated above, and a convolution operation is performed to finally obtain fine-grained features.

[0135] In specific implementation, the process of obtaining fine-grained features can be described by formula (8), as follows:

[0136]

[0137] In formula (8), Fdetail Represents fine-grained features of sample images; Represents the calculation of the complement of the mask feature, which is used to mask the low-frequency component and extract the high-frequency component; Used to indicate that the extracted k The high-frequency components covered by the frequency channels are extracted The high frequency components in It means multiplying the elements of the corresponding positions of the two matrices; DyConv The parentheses ( ) represent the dynamic convolution operation, an adaptive convolution operation that dynamically adjusts the weights of the convolution layer based on the input features. The other parameters have been explained above and will not be repeated here.

[0138] In the disclosed embodiment, the high-frequency component primarily contains detailed information from the sample image. By extracting the high-frequency component, detailed information closely related to thread fault detection can be extracted, providing more targeted data for subsequent processing. By determining the convolutional layer weights of the dynamic convolution layer based on the global information of the first global feature, the convolution operation can be made more targeted and adaptable, allowing the dynamic convolution layer to flexibly adjust the convolution operation based on the global situation of the sample image, improving the accuracy of image detail feature extraction, and thus further improving the detection accuracy of the thread fault detection model.

[0139] In the embodiment of the present disclosure, in order to integrate feature information at different levels to obtain more comprehensive and richer fusion features, the first global feature, the intermediate semantic feature and the fine-grained feature can be input into the fusion device of the model to be trained to obtain the fusion feature, such as Figure 6 As shown, it can be implemented as follows:

[0140] S601: Input the first global feature, the intermediate semantic feature, and the fine-grained feature into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature, respectively.

[0141] During the training phase, the aggregator includes a weight coefficient generator, which is built based on the attention mechanism and consists of a second global average pooling layer, a third global average pooling layer, a multilayer perceptron, and a third activation layer. The first global features, mid-level semantic features, and fine-grained features are input into the weight coefficient generator. Based on the attention mechanism, the weight coefficient generator calculates the importance weight of each feature. Specifically, by focusing on the contribution of each feature to the final classification result, it assigns higher weights to features with greater contributions and lower weights to features with lesser contributions, thereby obtaining the corresponding weight coefficients.

[0142] It's important to note that during the training phase, the weight coefficient generator dynamically generates weight coefficients based on the input features at different levels, allowing the trained model to automatically learn how to fuse features at different levels to optimize the objective function. After training, the trained model has learned the optimal weights for fusing features at different levels. Therefore, the weight coefficient generator branch can be removed after training, allowing the trained model to be simplified and optimized while maintaining performance, meeting the efficiency requirements of the subsequent inference phase.

[0143] During implementation, the first global feature, the intermediate semantic feature, and the fine-grained feature are input into a weight coefficient generator to obtain the weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature, respectively. This can be achieved based on the following steps:

[0144] Step B1, using a second global average pooling layer to process the intermediate semantic features to obtain a second global feature;

[0145] The mid-level semantic features are fed into the second global average pooling layer, which performs a global average pooling operation on the mid-level semantic features to generate the second global features. The second global features represent a global overview of the mid-level semantic features, which helps to grasp the mid-level semantic features from a holistic perspective when subsequently considering features at different levels.

[0146] Step B2: using a third global average pooling layer to process the fine-grained features to obtain a third global feature;

[0147] That is, the fine-grained features are input to the third global average pooling layer, which performs a global average pooling operation on the fine-grained features to obtain the third global feature. The third global feature reflects the overall characteristics of the fine-grained features.

[0148] Step B3, inputting the first global feature, the second global feature, and the third global feature into a multilayer perceptron to obtain a multilayer perceptron feature;

[0149] The first, second, and third global features are input into a multilayer perceptron. The multilayer perceptron, composed of multiple neural network layers, performs nonlinear transformations and fusion on the input features, ultimately outputting multilayer perceptual features. These features combine information from three levels of features and can further reflect the inherent connections between different features.

[0150] Step B4: Input the multi-layer perception features into the third activation layer to obtain the weight coefficient.

[0151] That is, the multi-layer perception features are input into the third activation layer, which performs nonlinear transformation through the activation function to obtain the weight coefficient.

[0152] In specific implementation, the process of obtaining the weight coefficient can be described by formula (9), as follows:

[0153]

[0154] In formula (9), are the weight coefficients corresponding to the first global feature, intermediate semantic feature and fine-grained feature respectively; MLP is represented by a multi-layer perceptron (i.e., multi-layer perceptron), which is used to map input features to output space. Softmax () function, which is used to convert the input vector into a probability distribution, that is, the weight coefficients corresponding to the first global feature, the intermediate semantic feature and the fine-grained feature in the implementation of the present disclosure, and the output value range is between [0,1] and the sum is 1. Indicates that a global average pooling operation is performed on the input; other parameters have been explained above and will not be repeated here.

[0155] In the embodiment of the present disclosure, the first global feature, the intermediate semantic feature and the fine-grained feature are input into the weight coefficient generator. The weight coefficient generator can learn the complex relationship between the global features at different levels based on the attention mechanism, so that when the model to be trained faces different spinning scenarios and silk thread fault types that may require different levels of feature combinations for detection, reasonable weight coefficients can be assigned to these different levels of features, so that the model to be trained can be more flexible in fusing features at different levels, thereby improving the accuracy of the final silk thread fault detection model in detecting silk thread faults.

[0156] S602: Perform weighted summation on the first global feature, the intermediate semantic feature, and the fine-grained feature based on the fuser to obtain a fused feature.

[0157] That is, the first global feature, the intermediate semantic feature, and the fine-grained feature are multiplied by their corresponding weight coefficients and then added together to obtain the fused feature.

[0158] In the specific implementation, fine-grained features are obtained, and the implementation process of the fusion device can be described by formula (10), as follows:

[0159]

[0160] In formula (10), Ffused represents fusion features; up () is an upsampling operation, which is used to adjust the size of the first global feature to the same size as the mid-level semantic features and fine-grained features; They represent the weight coefficients corresponding to the first global feature, intermediate semantic feature and fine-grained feature respectively.

[0161] In the disclosed embodiment, the first global feature, the intermediate semantic feature, and the fine-grained feature are input into a weight coefficient generator. The weight coefficient generator can adaptively learn and adjust the weights of each feature based on the input features, so that the trained model can more effectively utilize these features in different situations. The weighted summation of the first global feature, the intermediate semantic feature, and the fine-grained feature can integrate feature information of different scales and levels into a unified fusion feature, which can describe the characteristics of the sample image from different perspectives, avoid the limitations of a single feature, and thus improve the accuracy of wire fault detection.

[0162] In the embodiment of the present disclosure, in order to further improve the accuracy of the wire fault detection model in detecting wire faults, the model parameters of the to-be-trained model may be optimized based on the classification results, such as Figure 7 As shown, it can be implemented as follows:

[0163] S701: Determine the detection loss based on the classification result.

[0164] The detection loss includes at least one of the following:

[0165] (1) A global loss between the image to be optimized and the sample image generated based on the first global feature;

[0166] During implementation, a diffusion model may be used to process the first global feature, with the goal of restoring a sample image based on the first global feature to obtain an image to be optimized.

[0167] This global loss focuses on the differences at the overall image level and is used to measure the degree of closeness between the image to be optimized generated by the model and the sample image in terms of global features.

[0168] In the embodiment of the present disclosure, the global loss can be expressed as shown in formula (11):

[0169]

[0170] In formula (11), L1 represents the global loss; ||()||2 represents L 2 norm (Euclidean norm); F g represents the image to be optimized restored based on the first global feature; y Represents a sample image.

[0171] (2) Intermediate semantic loss; the intermediate semantic loss includes the accumulated value of the Hadamard product of each band mask and the corresponding residual in the mask feature; the residual includes the difference between the intermediate semantic feature and the target intermediate layer feature of the teacher model;

[0172] The mid-level semantic loss is used for knowledge distillation to transfer the knowledge of the intermediate-level features of the teacher model to the intermediate-level features of the current model.

[0173] In the embodiment of the present disclosure, the mid-level semantic loss can be expressed as shown in formula (12):

[0174]

[0175] In formula (12), L2 represents mid-level semantic loss; Mk is the mask feature; Fmid Represents mid-level semantic features; Represents the characteristics of the target middle layer of the teacher model; represents element-wise multiplication (Hadamard product); ||()||1 represents the L1 norm.

[0176] (3) Detail loss; Detail loss is determined based on the total variation regularization term of fine-grained features; The total variation regularization term is used to suppress high-frequency noise in fine-grained features;

[0177] That is, detail loss focuses more on the detail level of the sample image, which is used to reduce noise in fine-grained features while retaining edge information.

[0178] In the embodiment of the present disclosure, the mid-level semantic loss can be expressed as shown in formula (13):

[0179]

[0180] In formula (13), L 3 Indicates loss of detail; TV () represents the total variation regularization term, which is used to reduce noise and incoherence in detail features; F detail Represents fine-grained features.

[0181] (4) Classification loss: Classification loss represents the loss value between the classification result and the target classification result.

[0182] S702: Optimize model parameters of the to-be-trained model based on the detection loss.

[0183] During implementation, the calculated detection loss can be used to update the parameters of the model to be trained through the back-propagation algorithm, thereby reducing the loss value and improving the performance of the model to be trained.

[0184] In the disclosed embodiment, the detection loss measures the difference between the model output and the target from different aspects. By comprehensively considering the above-mentioned multiple detection losses, the model can be constrained and optimized from multiple dimensions such as the global information, local semantics, detailed features and final classification results of the sample image, thereby improving the accuracy of the model to be trained in the wire fault detection task.

[0185] In summary, the overall process of the wire fault detection model is as follows: Figure 8 As shown, including:

[0186] S801 , performing discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image, where the frequency domain representation includes a zero-frequency component, a low-frequency component, and a high-frequency component.

[0187] S802, input the sample image into the first global average pooling layer in the zero-frequency enhancer to obtain the global average pooling feature of the sample image; then, splice the zero-frequency component extracted from the frequency domain representation and the global average pooling feature in the channel dimension to obtain a spliced feature; finally, input the spliced feature into the convolution layer in the zero-frequency enhancer for convolution operation to obtain the first global feature.

[0188] S803, input the frequency domain representation into the first activation layer of the learnable frequency domain selector in the low-frequency reconstructor to obtain the activation feature; then input the obtained activation feature into the second activation layer of the learnable frequency domain selector to obtain the frequency weight matrix; then perform an inverse discrete cosine transform on the frequency weight matrix to convert the frequency weight matrix from the frequency domain back to the spatial domain to obtain the first spatial domain feature; then multiply the first spatial domain feature and the mask feature by element-by-element multiplication to obtain the intermediate feature; input the intermediate feature into the normalization layer to obtain the intermediate semantic feature.

[0189] S804, extracting high-frequency components from the frequency domain representation, converting the high-frequency components to the spatial domain, and obtaining second spatial domain features; inputting the second spatial domain features into the dynamic convolution layer in the high-frequency refiner to obtain fine-grained features.

[0190] S805 , performing weighted summation on the first global feature, the intermediate semantic feature, and the fine-grained feature based on the fusion device in the model to be trained to obtain a fusion feature.

[0191] S806: Input the fused features into a classifier to obtain a classification result for the sample image, wherein the classification result includes a floating yarn identification result and a broken yarn identification result.

[0192] S807: Based on the classification result, determine the detection losses such as global loss, mid-level semantic loss, detail loss, and classification loss between the image to be optimized generated based on the first global feature and the sample image, and optimize the parameters of the model to be trained based on the detection losses.

[0193] Based on the same concept, a method for detecting a thread fault in a spinning process is proposed in an embodiment of the present disclosure. The method is implemented based on the thread fault detection model in the spinning process described above. Figure 9 As shown, it can be implemented as:

[0194] S901, collecting images of the spinning manifold to obtain images to be processed.

[0195] During implementation, suitable image acquisition equipment, such as a high-speed industrial camera, can be used to capture images of the yarn at specific locations on the spinning beam to serve as input data for a subsequent yarn fault detection model.

[0196] S902 : Inputting the image to be processed into a yarn fault detection model obtained by training a yarn fault detection model in a spinning process to detect whether the yarn has problems of floating yarns and broken yarns.

[0197] The acquired image to be processed is input into a yarn fault detection model developed based on the training method for the yarn fault detection model in the spinning process described above. This yarn fault detection model extracts multi-level features from the input image, including global features, mid-level semantic features, and fine-grained features. It then fuses these features at different levels using methods such as a gated fusion mechanism to generate a fused feature. Finally, the model's classifier classifies the fused feature to determine whether the image to be processed has loose or broken yarns, and outputs a classification result based on the judgment.

[0198] The internal processing of the thread fault detection model is consistent with the previous description and will not be repeated here.

[0199] In the embodiments of the present disclosure, the provided method for detecting thread faults in the spinning process can realize automated and precise thread fault detection in the spinning process, which helps to improve production efficiency, reduce production costs, and improve product quality.

[0200] Based on the same technical concept, the present disclosure provides a training device 1000 for a yarn fault detection model in a spinning process, such as Figure 10 As shown, including:

[0201] The frequency domain decomposition module 1001 is used to perform discrete cosine transform on the sample image to obtain a frequency domain representation of the sample image; the sample image is obtained based on image acquisition of the spinning beam;

[0202] A zero-frequency enhancement module 1002 is configured to input the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of the model to be trained to obtain a first global feature of the sample image;

[0203] A low-frequency reconstruction module 1003 is used to input the frequency domain representation into the low-frequency reconstruction of the model to be trained to obtain the intermediate semantic features of the sample image;

[0204] A high-frequency refinement module 1004 is configured to input the frequency domain representation and the first global feature into a high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image;

[0205] A fusion module 1005 is configured to input the first global feature, the intermediate semantic feature, and the fine-grained feature into a fuser of a model to be trained to obtain a fused feature;

[0206] The classification module 1006 is used to input the fusion features into the classifier of the training model to obtain the classification results for the sample image, including the results of the floating silk discrimination and the broken silk discrimination;

[0207] The optimization module 1007 is used to optimize the model parameters of the to-be-trained model based on the classification result to obtain a wire fault detection model.

[0208] In some embodiments, wherein the zero-frequency enhancer comprises a first global average pooling layer and a convolutional layer;

[0209] Zero-frequency enhancement module, including:

[0210] A pooling unit, configured to input the sample image into a first global average pooling layer to obtain a global average pooling feature of the sample image;

[0211] A splicing unit is used to splice the zero-frequency component and the global average pooling feature to obtain a splicing feature;

[0212] The global extraction unit is used to input the concatenated features into the convolution layer to obtain the first global features.

[0213] In some embodiments, wherein the low frequency regenerator comprises a learnable frequency domain selector and a normalization layer;

[0214] Low frequency reconstruction module, including:

[0215] a weight learning unit, configured to learn a correlation pattern between different frequency bands in the frequency domain representation based on the learnable frequency domain selector, and generate a frequency weight matrix for selecting frequencies based on the correlation pattern;

[0216] an inverse transform unit, configured to perform an inverse discrete cosine transform on the frequency weight matrix to obtain a first spatial domain feature;

[0217] an intermediate feature determination unit, configured to determine an intermediate feature based on the first spatial domain feature and the mask feature; the mask feature including a frequency band mask corresponding to each frequency in the frequency domain representation;

[0218] The semantic feature determination unit is used to input the intermediate features into the normalization layer to obtain the intermediate semantic features.

[0219] In some embodiments, the weight learning unit is specifically configured to:

[0220] The frequency domain representation is input into the first activation layer of the learnable frequency domain selector to obtain activation features; the activation features are used to represent the correlation pattern between different frequency bands in the frequency domain representation;

[0221] The activated features are input into the second activation layer of the learnable frequency domain selector to obtain the frequency weight matrix.

[0222] In some embodiments, wherein the high frequency refiner comprises a dynamic convolutional layer;

[0223] High-frequency refinement module, including:

[0224] A high frequency extraction unit, used to extract high frequency components from the frequency domain representation;

[0225] A conversion unit, configured to convert the high-frequency component into a spatial domain to obtain a second spatial domain feature;

[0226] a weight determination unit, configured to determine a convolutional layer weight of the dynamic convolutional layer based on the first global feature;

[0227] The fine feature determination unit is used to input the second spatial domain feature into the dynamic convolution layer using the convolution layer weight to obtain fine-grained features.

[0228] In some embodiments, the fusion module includes:

[0229] A weight generation unit is used to input the first global feature, the intermediate semantic feature, and the fine-grained feature into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature respectively; the weight coefficient generator is constructed based on the attention mechanism;

[0230] The fusion unit is used to perform weighted summation on the first global feature, the intermediate semantic feature and the fine-grained feature based on the fuser to obtain a fused feature.

[0231] In some embodiments, wherein the weight coefficient generator includes a second global average pooling layer, a third global average pooling layer, a multilayer perceptron, and a third activation layer;

[0232] The weight generation unit is specifically used to:

[0233] The second global average pooling layer is used to process the intermediate semantic features to obtain the second global features;

[0234] The third global average pooling layer is used to process fine-grained features to obtain the third global features;

[0235] Inputting the first global feature, the second global feature and the third global feature into a multilayer perceptron to obtain a multilayer perceptron feature;

[0236] The multi-layer perception features are input into the third activation layer to obtain the weight coefficients.

[0237] In some embodiments, the optimization module includes:

[0238] a loss determination unit, configured to determine a detection loss based on a classification result;

[0239] An optimization unit, used to optimize the model parameters of the model to be trained based on the detection loss;

[0240] Detection loss includes at least one of the following:

[0241] A global loss between the image to be optimized and the sample image generated based on the first global feature;

[0242] Mid-level semantic loss; the mid-level semantic loss includes the accumulated value of the Hadamard product of each band mask and the corresponding residual in the mask feature; the residual includes the difference between the mid-level semantic feature and the feature map of the target intermediate layer of the teacher model;

[0243] Detail loss; Detail loss is determined based on the total variation regularization term of fine-grained features; The total variation regularization term is used to suppress high-frequency noise in fine-grained features;

[0244] Classification loss: Classification loss represents the loss value between the classification result and the target classification result.

[0245] Based on the same technical concept, the present disclosure provides a device 1100 for detecting yarn faults in a spinning process. Figure 11 As shown, including:

[0246] The acquisition module 1101 is used to acquire images of the spinning beam to obtain images to be processed;

[0247] The detection module 1102 is used to input the image to be processed into a yarn fault detection model obtained by a training method based on a yarn fault detection model in a spinning process, so as to detect whether the yarn has floating yarn or broken yarn problems.

[0248] For the description of specific functions and examples of each module, sub-module\unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0249] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0250] Figure 12 FIG. 1 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 12 As shown, the electronic device includes: a memory 1210 and a processor 1220. The memory 1210 stores a computer program that can be executed on the processor 1220. The number of memories 1210 and processors 1220 can be one or more. The memory 1210 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device performs the method provided by the above method embodiment. The electronic device may also include: a communication interface 1230 for communicating with external devices and performing data exchange.

[0251] If the memory 1210, the processor 1220, and the communication interface 1230 are implemented independently, the memory 1210, the processor 1220, and the communication interface 1230 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0252] Optionally, in a specific implementation, if the memory 1210, the processor 1220 and the communication interface 1230 are integrated on a chip, the memory 1210, the processor 1220 and the communication interface 1230 can communicate with each other through an internal interface.

[0253] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0254] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).

[0255] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, data subscriber line (DSL)) or wireless (e.g., infrared, Bluetooth, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)). It is worth noting that the computer-readable storage medium mentioned in the present disclosure may be a non-volatile storage medium, in other words, a non-transient storage medium.

[0256] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0257] In the description of the embodiments of the present disclosure, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0258] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means or. For example, A / B can mean A or B. "And / or" in this document is only a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0259] In the description of the embodiments of the present disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.

[0260] The above description is merely an exemplary embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.

Claims

1. A method for training a yarn fault detection model in a spinning process, characterized in that: include: Performing discrete cosine transform on the sample image to obtain a frequency domain representation of the sample image; The sample image is obtained based on image acquisition of the spinning beam; Inputting the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a to-be-trained model to obtain a first global feature of the sample image; Inputting the frequency domain representation into the low-frequency reconstructor of the to-be-trained model to obtain the intermediate semantic features of the sample image; Inputting the frequency domain representation and the first global feature into a high-frequency refiner of the to-be-trained model to obtain fine-grained features of the sample image; Inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into the fusion device of the to-be-trained model to obtain a fused feature; Inputting the fusion features into the classifier of the model to be trained to obtain a classification result for the sample image, wherein the classification result includes a floating silk discrimination result and a broken silk discrimination result; Based on the classification result, the model parameters of the model to be trained are optimized to obtain a wire fault detection model.

2. The method according to claim 1, characterized in that The zero-frequency enhancer includes a first global average pooling layer and a convolutional layer; Inputting the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a to-be-trained model to obtain a first global feature of the sample image includes: Inputting the sample image into the first global average pooling layer to obtain a global average pooling feature of the sample image; Splicing the zero-frequency component and the global average pooling feature to obtain a splicing feature; The concatenated features are input into the convolutional layer to obtain the first global features.

3. The method according to claim 1, characterized in that The low frequency regenerator includes a learnable frequency domain selector and a normalization layer; The step of inputting the frequency domain representation into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image comprises: learning a correlation pattern between different frequency bands in the frequency domain representation based on the learnable frequency domain selector, and generating a frequency weight matrix for selecting frequencies based on the correlation pattern; Performing an inverse discrete cosine transform on the frequency weight matrix to obtain a first spatial domain feature; Determining an intermediate feature based on the first spatial domain feature and the mask feature; the mask feature includes a frequency band mask corresponding to each frequency in the frequency domain representation; The intermediate features are input into the normalization layer to obtain the intermediate semantic features.

4. The method according to claim 3, characterized in that The learning of the correlation pattern between different frequency bands in the frequency domain representation based on the learnable frequency domain selector and generating a frequency weight matrix for selecting frequencies based on the correlation pattern includes: Inputting the frequency domain representation into the first activation layer of the learnable frequency domain selector to obtain activation features; the activation features are used to represent the correlation pattern between different frequency bands in the frequency domain representation; The activation features are input into the second activation layer of the learnable frequency domain selector to obtain the frequency weight matrix.

5. The method according to any one of claims 1 to 4, characterized in that The high-frequency refiner includes a dynamic convolution layer; Inputting the frequency domain representation and the first global feature into the high-frequency refiner of the to-be-trained model to obtain the fine-grained features of the sample image includes: extracting high frequency components from the frequency domain representation; Converting the high-frequency component into a spatial domain to obtain a second spatial domain feature; Determining a convolutional layer weight of the dynamic convolutional layer based on the first global feature; The second spatial domain feature is input into the dynamic convolution layer using the convolution layer weight to obtain the fine-grained feature.

6. The method according to claim 1, characterized in that The step of inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into the fusion device of the to-be-trained model to obtain a fused feature includes: Inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature, respectively; the weight coefficient generator is constructed based on an attention mechanism; Based on the fuser, a weighted sum is performed on the first global feature, the mid-level semantic feature, and the fine-grained feature to obtain the fused feature.

7. The method according to claim 6, characterized in that The weight coefficient generator includes a second global average pooling layer, a third global average pooling layer, a multilayer perceptron and a third activation layer; The step of inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature, respectively, includes: Processing the intermediate semantic features using the second global average pooling layer to obtain a second global feature; Processing the fine-grained features using the third global average pooling layer to obtain a third global feature; Inputting the first global feature, the second global feature, and the third global feature into the multilayer perceptron to obtain a multilayer perceptron feature; The multi-layer perception features are input into the third activation layer to obtain the weight coefficient.

8. The method according to claim 1, characterized in that Optimizing the model parameters of the to-be-trained model based on the classification result includes: determining a detection loss based on the classification result; Optimizing model parameters of the to-be-trained model based on the detection loss; The detection loss includes at least one of the following: A global loss between the image to be optimized and the sample image generated based on the first global feature; A mid-level semantic loss; the mid-level semantic loss comprises the cumulative value of the Hadamard product of each band mask and the corresponding residual in the mask feature; the residual comprises the difference between the mid-level semantic feature and the feature map of the target intermediate layer of the teacher model; Detail loss; the detail loss is determined based on a total variation regularization term of the fine-grained features; the total variation regularization term is used to suppress high-frequency noise in the fine-grained features; Classification loss; the classification loss represents the loss value between the classification result and the target classification result.

9. A method for detecting yarn faults in a spinning process, characterized in that: include: Capturing images of the spinning beam to obtain images to be processed; The image to be processed is input into a silk thread fault detection model obtained by the method according to any one of claims 1 to 8 to detect whether the silk thread has problems of floating or broken threads.

10. A training device for a yarn fault detection model in a spinning process, characterized in that: include: A frequency domain decomposition module, configured to perform discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; The sample image is obtained based on image acquisition of the spinning beam; a zero-frequency enhancement module, configured to input the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a to-be-trained model to obtain a first global feature of the sample image; a low-frequency reconstruction module, configured to input the frequency domain representation into a low-frequency reconstruction device of the to-be-trained model to obtain intermediate semantic features of the sample image; a high-frequency refinement module, configured to input the frequency domain representation and the first global feature into a high-frequency refiner of the to-be-trained model to obtain fine-grained features of the sample image; A fusion module, configured to input the first global feature, the intermediate semantic feature, and the fine-grained feature into a fuser of the to-be-trained model to obtain a fused feature; a classification module, configured to input the fusion features into the classifier of the model to be trained to obtain a classification result for the sample image, wherein the classification result includes a floating silk discrimination result and a broken silk discrimination result; An optimization module is used to optimize the model parameters of the to-be-trained model based on the classification result to obtain a wire fault detection model.

11. A yarn fault detection device in a spinning process, characterized in that: include: An acquisition module is used to acquire images of the spinning box to obtain images to be processed; A detection module is used to input the image to be processed into a silk thread fault detection model obtained by the device according to claim 10 to detect whether the silk thread has problems of floating silk and broken silk.

12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

14. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Two-dimensional feature selection verification method for optical remote sensing image similarity evaluation

    CN118506135A

  • Remote sensing image segmentation method based on channel enhancement and cross-level multi-input features

    CN119380018A