Training method of silk thread fault detection model in spinning process and related device
Through deep learning technology, discrete cosine transformation and multi-layer feature extraction methods are used to solve the accuracy of wire fault detection in spinning process, efficient wire fault detection is achieved, and the automation and product quality of spinning production are improved.
Patent Information
- Application Number
- CN202510783720.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-12
AI Technical Summary
In spinning processes, it is difficult for the prior art to efficiently and accurately detect wire failures, such as broken wires and floating wires, resulting in low production efficiency and unstable product quality.
The wire fault detection model based on deep learning is adopted, and the frequency domain representation is obtained through discrete cosine transformation. Multi-level features are extracted by combining the zero-frequency enhancer, low-frequency reconstructor and high-frequency refiner, and fault judgment is performed through the fusion and classifier to optimize the model parameters to improve detection accuracy.
Accurate detection of wire failures has been achieved, the automation level of spinning production has been improved, production costs have been reduced, product quality and production efficiency have been improved.
Smart Images

Figure CN120279346A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and in particular to technologies such as spinning technology and deep learning. Background Art
[0002] In the industrial scenario of the spinning process, the quality of the silk thread directly determines the performance and quality of the final product. As the core output of the spinning process, the quality of the silk thread plays a decisive role in key performance indicators such as the strength, toughness, and dyeing consistency of downstream textiles. Summary of the Invention
[0003] The present disclosure provides a method for training a silk thread fault detection model in a spinning process to solve or alleviate one or more technical problems in the related art.
[0004] In a first aspect, the present disclosure provides a method for training a silk thread fault detection model in a spinning process, including: Performing discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; the sample image is obtained based on image acquisition of a spinning box; Inputting the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a model to be trained to obtain a first global feature of the sample image; Inputting the frequency domain representation into a low-frequency reconstructor of the model to be trained to obtain intermediate semantic features of the sample image; Inputting the frequency domain representation and the first global feature into a high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image; Inputting the first global feature, the intermediate semantic features, and the fine-grained features into a fuser of the model to be trained to obtain a fused feature; Inputting the fused feature into a classifier of the model to be trained to obtain a classification result for the sample image, where the classification result includes a floating silk discrimination result and a broken silk discrimination result; Optimizing the model parameters of the model to be trained based on the classification result to obtain a silk thread fault detection model.
[0005] In a second aspect, the present disclosure provides a training device for a silk thread fault detection model in a spinning process, including: A frequency domain decomposition module for performing discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; the sample image is obtained based on image acquisition of a spinning box; A zero-frequency enhancement module for inputting the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a model to be trained to obtain a first global feature of the sample image; A low-frequency reconstruction module for inputting the frequency domain representation into a low-frequency reconstructor of the model to be trained to obtain intermediate semantic features of the sample image; A high-frequency refinement module for inputting the frequency-domain representation and the first global feature into a high-frequency refiner of a model to be trained to obtain fine-grained features of a sample image; A fusion module for inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into a fuser of a model to be trained to obtain a fused feature; A classification module for inputting the fused feature into a classifier of a model to be trained to obtain a classification result for the sample image, where the classification result includes a floating wire discrimination result and a broken wire discrimination result; An optimization module for optimizing model parameters of a model to be trained based on the classification result to obtain a wire fault detection model.
[0006] In a third aspect, the present disclosure provides a method for detecting wire faults in a spinning process, including: Performing image acquisition on a spinning box to obtain an image to be processed; Inputting the image to be processed into a wire fault detection model obtained by a training method of a wire fault detection model in a spinning process to detect whether there are floating wire and broken wire problems in the wire.
[0007] In a fourth aspect, the present disclosure provides a device for detecting wire faults in a spinning process, including: An acquisition module for performing image acquisition on a spinning box to obtain an image to be processed; A detection module for inputting the image to be processed into a wire fault detection model obtained by a training method of a wire fault detection model in a spinning process to detect whether there are floating wire and broken wire problems in the wire.
[0008] In a fifth aspect, there is provided an electronic device, including: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0009] In a sixth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0010] In a seventh aspect, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements any method in the embodiments of the present disclosure.
[0011] In the embodiments of the present disclosure, the advantages of each feature can be utilized to enable the model to be trained to obtain a relatively comprehensive and rich image feature representation, so as to more accurately describe the sample image and provide more powerful feature support for subsequent classification. Based on the classification results, the parameters of the model to be trained are continuously optimized, enabling it to gradually learn how to better extract features and perform classification, so that the finally obtained silk thread fault detection model has a high classification accuracy on the training data set.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0014] Figure 1 is a schematic flowchart of a method for training a silk thread fault detection model in a spinning process according to an embodiment of the present disclosure; Figure 2 is a schematic flowchart of obtaining the first global feature of a sample image according to an embodiment of the present disclosure; Figure 3 is a schematic flowchart of obtaining the intermediate semantic feature of a sample image according to an embodiment of the present disclosure; Figure 4 is an exemplary mask feature map according to an embodiment of the present disclosure; Figure 5 is a schematic flowchart of obtaining the fine-grained feature of a sample image according to an embodiment of the present disclosure; Figure 6 is a schematic flowchart of obtaining the fused feature according to an embodiment of the present disclosure; Figure 7 is a schematic flowchart of optimizing the model parameters of the model to be trained according to an embodiment of the present disclosure; Figure 8 is an overall structural diagram of a method for training a silk thread fault detection model in a spinning process according to an embodiment of the present disclosure; Figure 9 is a schematic flowchart of a silk thread fault detection method in a spinning process according to an embodiment of the present disclosure; Figure 10 is a schematic structural diagram of a device for training a silk thread fault detection model in a spinning process according to an embodiment of the present disclosure; Figure 11 is a schematic structural diagram of a yarn fault detection device in a spinning process according to an embodiment of the present disclosure; Figure 12 It is a block diagram of an electronic device used to implement the training method of the yarn fault detection model in the spinning process and / or the yarn fault detection method in the spinning process according to the embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] The training method of the yarn fault detection model in the spinning process provided by the present disclosure will be further described in detail with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0016] In addition, in order to better illustrate the technical solutions of the embodiments of the present disclosure, numerous specific details are given in the specific implementations below. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components, circuits, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.
[0017] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0018] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order of multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.
[0019] As the chemical fiber industry's requirements for product quality continue to increase and production scale continues to expand, quality control in the silk thread production process becomes increasingly critical.
[0020] In the actual spinning production process, the silk thread will inevitably have abnormal phenomena such as broken silk and floating silk. Broken silk refers to the phenomenon that the silk thread suddenly breaks during the spinning process, which will affect the continuity and production efficiency of spinning. If it causes production interruption and generates a large amount of waste silk, it will bring great losses to the company; floating silk refers to the phenomenon that the silk thread floats irregularly during the spinning process. The floating silk phenomenon will cause problems with product quality and will also bring direct losses to the company. Once these abnormal situations occur, they will bring some negative effects.
[0021] Therefore, there is an urgent need for an efficient, accurate, and adaptable silk thread fault detection technology that can promptly and precisely detect abnormal conditions such as broken filaments and floating filaments, so as to ensure the quality of spun products, improve production efficiency, reduce production costs, and meet the growing production demands of the modern textile industry.
[0022] In view of this, in the embodiments of the present disclosure, a training method for a silk thread fault detection model in a spinning process and a silk thread fault detection method in the spinning process are proposed based on artificial intelligence technology.
[0023] It should be noted that the main types of spun products involved in the solutions of the embodiments of the present disclosure may include one or more of partially oriented yarns (POY), fully drawn yarns (FDY), polyester staple fiber, etc. For example, the types of silk can specifically include polyester partially oriented yarns, polyester fully drawn yarns, polyester drawn yarns, polyester staple fiber, etc.
[0024] The training method for the silk thread fault detection model and the silk thread fault detection method will be described separately below.
[0025] As Figure 1 shown, it is a schematic diagram of the processing flow of the training method for the silk thread fault detection model provided by the embodiments of the present disclosure, which can be specifically implemented as: S101, perform discrete cosine transform on the sample image to obtain the frequency domain representation of the sample image.
[0026] The sample image is obtained by collecting images of the spinning box and can represent the state of the silk thread at different times. After obtaining the sample image, in order to improve the model training efficiency and model detection results, the format of the sample image can be standardized first and adjusted to a unified size. Subsequently, the processed sample image is divided into pixel blocks of n×n, so as to provide data support for the subsequent discrete cosine transform (DCT), and to ensure the standardization and consistency of the processing process. Among them, the value of n is an integer greater than 1. For example, the value is 8, and the specific value can be determined according to the actual needs of the business and the needs of the DCT transform.
[0027] The discrete cosine transform is a transform technique related to the Fourier transform, which is used to transform a signal from one domain to the frequency domain. In the embodiments of the present disclosure, by performing the discrete cosine transform on the sample image, the sample image can be transformed from the spatial domain to different levels of the frequency domain, so as to obtain the frequency domain representation of the sample image. For example, the zero-frequency component in the frequency domain representation can capture the global information such as the overall brightness and contrast of the sample image; the low-frequency components in the frequency domain representation can reflect the regions with slow changes in the sample image, such as the parts where the silk threads are evenly distributed in the sample image; the high-frequency components in the frequency domain representation reflect the details with rapid changes in the sample image, such as the edge contours of the silk threads and the texture of the intertwined silk threads in the sample image.
[0028] Specifically, when implementing, the frequency domain representation of the sample image can be obtained through formula (1), which is as follows: In formula (1), represents the frequency domain representation of the sample image, that is, the DCT coefficient of the c th channel, the h th block row, and the w th block column; represents the pixel value of the x th row, the c th channel, and the nh+i th column in the sample image nw+j ; is the cosine basis function in the two-dimensional DCT transform, which is used to convert the pixel value in the spatial domain into the frequency domain coefficient, and this cosine basis function is used to calculate the frequency component in the horizontal direction; is used to calculate the frequency component in the vertical direction, u, v is the frequency index, ranging from 0 to n - 1 respectively; i, j is the pixel index within the pixel block, ranging from 0 to n - 1 respectively.
[0029] S102: Input the zero-frequency component in the frequency domain representation and the sample image into the zero-frequency enhancer of the model to be trained, and obtain the first global feature of the sample image.
[0030] In the frequency domain representation, the zero-frequency component usually refers to the frequency domain component corresponding to the case where the frequency index takes the value of u = 0 and v = 0. Thus, the zero-frequency component can be obtained based on formula (2): In formula (2), represents the zero-frequency component, and other parameters are the same as those in formula (1), which will not be elaborated here.
[0031] Input the zero-frequency component in the frequency-domain representation and the sample image into the zero-frequency enhancer of the model to be trained. The zero-frequency enhancer can generate a first global feature containing the global information of multiple dimensions of the sample image based on the overall information of the sample image provided by the zero-frequency component and the information provided by the sample image. That is to say, the zero-frequency enhancer is the first neural network module of the model to be trained, and is used to extract the first global feature from the zero-frequency component and the sample image.
[0032] S103. Input the frequency-domain representation into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image.
[0033] Among them, the low-frequency reconstructor is the second neural network module of the model to be trained, and is used to extract the features of the low-frequency component from the frequency-domain representation to obtain the intermediate semantic features.
[0034] It can also be understood that inputting the frequency-domain representation into the low-frequency reconstructor of the model to be trained, the low-frequency reconstructor will process the low-frequency component in the frequency-domain representation and extract the intermediate semantic features of the sample image. The intermediate semantic features help the model to be trained to understand the continuity features of the silk thread in the sample image, such as the part where the silk thread distribution in the sample image is relatively uniform.
[0035] S104. Input the frequency-domain representation and the first global feature into the high-frequency refiner of the model to be trained to obtain the fine-grained features of the sample image.
[0036] Among them, the high-frequency refiner is the third neural network module of the model to be trained, and is used to extract the features of the high-frequency component from the frequency-domain representation and the first global feature to obtain the detailed features of the high-frequency component.
[0037] That is, input the frequency-domain representation and the first global feature into the high-frequency refiner of the model to be trained. The high-frequency refiner will combine the high-frequency component in the frequency-domain representation and the first global feature to further extract the detailed information of the sample image. The fine-grained features can help the model to be trained to capture the high-frequency vibration trajectory features of the silk thread, such as the jitter edge of the silk thread and the specific form of motion blur.
[0038] S105. Input the first global feature, the intermediate semantic features and the fine-grained features into the fuser of the model to be trained to obtain the fused feature.
[0039] Input the first global feature, the intermediate semantic features and the fine-grained features into the fuser of the model to be trained. The fuser can combine the first global feature, the intermediate semantic features and the fine-grained features, make full use of the advantages of each level of features, and generate a fused feature with a multi-semantic level feature representation.
[0040] S106. Input the fused feature into the classifier of the model to be trained to obtain the classification result for the sample image. The classification result includes the floating silk discrimination result and the broken silk discrimination result.
[0041] The classifier can analyze the fused features to determine whether there is a floating wire phenomenon or a broken wire phenomenon in the sample image, and output corresponding discrimination results based on this, that is, clearly indicate whether the wire in the current sample image is a floating wire, a broken wire, or in a normal state, etc.
[0042] S107. Based on the classification results, optimize the model parameters of the model to be trained to obtain a wire fault detection model.
[0043] During implementation, based on the classification results obtained by the above classifier, compare with the corresponding true labels, calculate metrics such as loss values, and then optimize the model parameters of the model to be trained accordingly, so that the model to be trained can continuously improve the accuracy of the inference process, and finally obtain an effective wire fault detection model for wire fault detection of newly acquired sample images in the actual spinning process.
[0044] In the embodiments of the present disclosure, perform a discrete cosine transform on the sample image to obtain a frequency-domain representation of the sample image, which can convert the sample image collected based on the spinning box from the spatial domain to the frequency domain, so as to facilitate subsequent targeted processing of different features corresponding to different frequency components. Input the zero-frequency component in the frequency-domain representation and the sample image into the zero-frequency enhancer of the model to be trained to obtain the first global feature of the sample image, which can help the model to be trained extract the features of the sample image as a whole. Input the frequency-domain representation into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic feature of the sample image, which can help the model to be trained understand the continuity feature of the wire in the sample image, thereby providing a basis for detecting wire faults. Input the frequency-domain representation and the first global feature into the high-frequency refiner of the model to be trained to obtain the fine-grained feature of the sample image, which can help the model to be trained capture more subtle changes in the wire in the sample image. Input the first global feature, the intermediate semantic feature, and the fine-grained feature into the fuser of the model to be trained to obtain the fused feature. By fusing features of different dimensions and levels, the advantages of each feature can be fully utilized, enabling the model to be trained to obtain a relatively comprehensive and rich image feature representation, thereby more accurately describing the sample image and providing stronger feature support for subsequent classification. Continuously optimize the model parameters of the model to be trained based on the classification results, so that it can gradually learn how to better extract features and perform classification, so that the finally obtained wire fault detection model has a high classification accuracy on the training data set.
[0045] In the embodiments of the present disclosure, the zero-frequency enhancer includes a first global average pooling layer and a convolutional layer.
[0046] The global average pooling layer is a pooling operation commonly used in deep learning. It can perform global average pooling on the input feature map, taking the average of all elements in each feature channel to obtain a value that represents the overall information of the feature channel. In the embodiments of the present disclosure, the first global average pooling layer receives the sample image as input and outputs the global average pooling feature of the sample image.
[0047] The convolutional layer is an important structure for feature extraction in deep learning. By sliding the convolutional kernel over the input data and performing convolutional operations, it can automatically learn local features and patterns in the data. In the embodiments of the present disclosure, the convolutional layer receives the concatenated feature as input and outputs the first global feature.
[0048] During implementation, the zero-frequency component in the frequency-domain representation and the sample image are input into the zero-frequency enhancer of the model to be trained, and the first global feature of the sample image is obtained, as Figure 2 shown, and it can be specifically implemented as: S201, input the sample image into the first global average pooling layer to obtain the global average pooling feature of the sample image.
[0049] The global average pooling feature is obtained by taking the average of each channel of the sample image in the spatial dimension (usually height and width), and its calculation method is shown in formula (3): In formula (3), represents the global average pooling feature of the sample image; represents the sample image x in the channel dimension c , and the elements in the spatial dimension i (height direction) and j (width direction); h and w respectively represent the sizes of the tensor in the height and width dimensions.
[0050] S202, concatenate the zero-frequency component and the global average pooling feature to obtain the concatenated feature.
[0051] During implementation, the zero-frequency component in the frequency-domain representation and the global average pooling feature obtained in the previous step can be concatenated in the channel dimension to obtain the concatenated feature.
[0052] S203, input the concatenated feature into the convolutional layer to obtain the first global feature.
[0053] Input the concatenated feature into the convolutional layer for convolutional operations, that is, by fusing and reducing the dimensions between channels of the concatenated feature, the final first global feature is generated.
[0054] During implementation, the acquisition process of the first global feature can be expressed by formula (4): In formula (4), represents the global average pooling feature of the sample image; represents the zero-frequency component in the frequency-domain representation; represents the global average pooling feature of the sample image; represents the concatenation operation in the channel dimension, which is used to concatenate the zero-frequency component in the frequency-domain representation and the global average pooling feature of the sample image; represents a 1×1 convolution operation, which is used to extract the first global feature of the sample image without changing the spatial size of the feature map.
[0055] In the embodiments of the present disclosure, through global average pooling, the feature maps on each channel of the sample image can be averaged, and the feature maps on each channel are converted into a scalar value, reducing the dimension of the features, thereby helping the model to be trained to reduce the computational complexity of the model. Concatenating the zero-frequency component and the global average pooling feature to obtain a concatenated feature can integrate the brightness information and the global feature of the sample image, enabling the model to be trained to consider both the brightness and the overall feature of the sample image simultaneously, and improving the model's understanding ability of the sample image. Inputting the concatenated feature into the convolutional layer, thereby fusing the global brightness and the first global feature and representing them in a higher-level feature space, is conducive to obtaining the first global feature that can better reflect the overall feature of the sample image, and can provide a more effective feature representation for subsequent fault classification, thereby improving the accuracy of wire fault detection.
[0056] In the embodiments of the present disclosure, the low-frequency reconstructor includes a learnable frequency-domain selector and a normalization layer.
[0057] The learnable frequency-domain selector learns the association patterns between different frequency bands in the frequency-domain representation, thereby generating a frequency weight matrix for selecting frequencies, and can automatically determine which frequency band features are more helpful for extracting reasonable intermediate semantic features. The learnable frequency-domain selector is a key component for screening and integrating frequency-domain information, and through training optimization, it can learn the classification and recognition pattern of wire faults and improve the accuracy of wire fault detection.
[0058] The normalization layer normalizes the intermediate features to make their data distribution more stable, which is helpful for subsequent feature extraction and model training processes, can eliminate the dimension and numerical range differences between different features, and accelerate the convergence speed of the model.
[0059] During implementation, input the frequency-domain representation into the low-frequency reconstructor of the model to be trained to obtain the intermediate semantic features of the sample image, as Figure 3 shown, and specifically, it can be implemented as: S301. Learn the association pattern between different frequency bands in the frequency domain representation based on a learnable frequency domain selector, and generate a frequency weight matrix for frequency selection based on the association pattern.
[0060] Input the frequency domain representation into the learnable frequency domain selector, which can analyze the association pattern between different frequency bands in the frequency domain representation through a learning algorithm. For example, in the frequency domain representation of a silk thread sample image, the low-frequency band may be related to the continuity feature of the silk thread, while the high-frequency band may be related to the local texture. The learnable frequency domain selector will calculate the importance weight of each frequency component according to these association patterns and generate a frequency weight matrix.
[0061] When implemented, generating a frequency weight matrix for frequency selection based on the association pattern can be achieved through the following steps: Step A1. Input the frequency domain representation into the first activation layer of the learnable frequency domain selector to obtain activation features; the activation features are used to represent the association pattern between different frequency bands in the frequency domain representation. That is, input the frequency domain representation into the first activation layer of the learnable frequency domain selector. The first activation layer performs a non-linear transformation on the frequency domain representation through an activation function to obtain activation features. The activation features can reflect the association pattern between different frequency bands in the frequency domain representation.
[0062] Step A2. Input the activation features into the second activation layer of the learnable frequency domain selector to obtain a frequency weight matrix.
[0063] Input the obtained activation features into the second activation layer of the learnable frequency domain selector. The second activation layer performs a non-linear transformation again through the activation function and finally obtains a frequency weight matrix. Each element in the frequency weight matrix corresponds to a frequency component in the frequency domain representation, and the value in the frequency weight matrix represents the importance weight of this frequency component in constructing the intermediate semantic features. The frequency components with higher weights will play a greater role in generating the intermediate semantic features, while the frequency components with lower weights may be suppressed or ignored.
[0064] In the embodiments of the present disclosure, the learnable frequency domain selector can be expressed as shown in formula (5): In formula (5), represents the first activation layer, W1 is the weight of the first activation layer, b1 is the bias of the first activation layer. The first activation layer projects the input frequency domain representation onto a new feature space, and then learns W1 the parameters of to capture the association pattern between different frequency bands; b1The bias term is used to adjust the distribution of features and helps the model better fit the data; is a non - linear activation function that can introduce non - linear characteristics, enabling the model to learn complex feature relationships. represents the second activation layer, which is used to generate a frequency - selection weight matrix by learning the parameters of the weights W2 on the basis of the activation features obtained from the first activation layer. This matrix is used to weight different frequency components; the bias term of the second activation layer b2 is used to adjust the distribution of the features in the second layer.
[0065] In the embodiments of the present disclosure, the learnable frequency - domain selector first inputs the frequency - domain representation into the first activation layer, obtains the correlation patterns between different frequency bands in the frequency - domain representation through non - linear transformation, and obtains the activation features representing these correlation patterns; then it inputs the activation features into the second activation layer to generate a frequency - weight matrix for frequency selection. This matrix can adaptively assign corresponding weights to different frequencies, so that the model to be trained can adaptively select the frequencies that are most helpful for detecting filament faults to adapt to different spinning scenarios and fault types. Compared with the fixed frequency - selection method, this adaptive method is more flexible and effective, and can improve the generalization ability of the model to be trained.
[0066] In the training stage, by continuously optimizing the weights and biases of the first and second activation layers in the learnable frequency - domain selector, the learnable frequency - domain selector can learn to capture important low - frequency component information. After the training is completed, the weights and biases of the first and second activation layers no longer change, can adapt to the frequency - domain representation, and screen out appropriate low - frequency components for constructing intermediate semantic features.
[0067] S302, perform an inverse discrete cosine transform on the frequency - weight matrix to obtain the first spatial - domain feature.
[0068] That is, perform an inverse discrete cosine transform on the generated frequency - weight matrix to convert it from the frequency domain back to the spatial domain to obtain the first spatial - domain feature.
[0069] S303, determine the intermediate feature based on the spatial - domain feature and the mask feature; the mask feature includes the band masks corresponding to each frequency in the frequency - domain representation.
[0070] Based on the above - obtained first spatial - domain feature and the pre - set mask feature, determine the intermediate feature. The mask feature contains the band masks corresponding to each frequency in the frequency - domain representation, such as Figure 4As shown, an exemplary mask matrix (i.e., mask feature) provided by the implementation of the present disclosure, where the upper left region can be used as a band mask corresponding to low frequencies to capture macroscopic features such as the silk thread contour in the sample image; the middle region can be used as a band mask corresponding to medium frequencies to extract the scale texture information of the silk thread in the sample image; the lower left region can be used as a band mask corresponding to high frequencies to retain more fine image detail information in the sample image. Through the band mask, the frequency components can be further refined to screen, and it is clear which frequency components should be retained or highlighted and which should be suppressed or discarded. For example, the first k×k low-frequency components in the frequency domain representation are retained, while other high-frequency components are masked out.
[0071] S304, input the intermediate feature into the normalization layer to obtain the intermediate semantic feature.
[0072] Input the determined intermediate feature into the normalization layer for normalization processing. The normalization layer will perform a normalization operation on the data of the intermediate feature to make it meet specific distribution requirements, thereby obtaining the intermediate semantic feature.
[0073] In specific implementation, the process of obtaining the intermediate semantic feature can be described by formula (6) as follows: In formula (6), Fmid represents the intermediate semantic feature of the sample image; Mk represents the mask feature of the first k × k low-frequency components in the frequency domain representation; k is a positive integer used to control the range of retained low-frequency components; represents the obtained frequency domain representation; represents the frequency weight matrix generated based on the association pattern for selecting frequencies; represents the inverse discrete cosine transform, which is used to convert the frequency weight matrix from the frequency domain representation back to the spatial domain feature; represents element-wise multiplication (Hadamard product), that is, multiplying the elements at the corresponding positions of the two matrices; LayerNorm () represents the normalization layer, which is used to process the intermediate feature to obtain the intermediate semantic feature.
[0074] In the embodiments of the present disclosure, the learnable frequency domain selector can deeply explore the complex correlation patterns between different frequency bands in the frequency domain representation. Based on the learnable frequency domain selector learning the correlation patterns and generating a frequency weight matrix, the matrix can adaptively assign corresponding weights to each frequency, so that the model can adaptively select the frequencies that are most helpful for detecting wire faults. Determine the intermediate features based on the spatial domain features and the mask features, so that the intermediate features reflect both the important features learned according to the frequency correlation in the spatial domain and meet the frequency screening requirements specified by the mask features. Normalize the intermediate features through a normalization layer, which helps to improve the training stability and convergence speed of the model to be trained, and avoid the training difficulty problem caused by too large differences in feature scales. At the same time, the normalized intermediate semantic features are more conducive to subsequent model processing and analysis, and improve the model's ability to detect wire faults in sample images.
[0075] In the embodiments of the present disclosure, the high-frequency refiner includes a dynamic convolution layer. Different from the traditional convolution layer, the convolution kernel weights of the dynamic convolution layer are not fixed, but can be dynamically generated according to the input features. This enables the model to adaptively adjust the convolution operation according to the characteristics of different sample images, so as to more accurately extract fine-grained features.
[0076] During implementation, input the frequency domain representation and the first global feature into the high-frequency refiner of the model to be trained to obtain the fine-grained features of the sample image, as Figure 5 shown, which can be specifically implemented as: S501, extract the high-frequency components from the frequency domain representation.
[0077] Based on the content described in the previous step S303, where the mask feature retains the first × low-frequency components in the frequency domain representation and masks other high-frequency components. Thus, by selecting the complement of the mask feature, the low-frequency components can be masked and only the high-frequency components are retained, so as to extract the high-frequency components from the frequency domain representation.
[0078] S502, convert the high-frequency components to the spatial domain to obtain the second spatial domain feature.
[0079] That is, convert the extracted high-frequency components to the spatial domain through the inverse discrete cosine transform to obtain the second spatial domain feature.
[0080] S503, determine the convolution layer weights of the dynamic convolution layer based on the first global feature.
[0081] The first global feature contains the overall information of the sample image. By using this global information to dynamically generate the weights of the convolutional layer, the convolution operation can be made more targeted. For example, if the first global feature indicates that there may be a silk thread fault feature in a certain area of the sample image, the generated weights of the convolutional layer may pay more attention to the detail extraction of this area.
[0082] That is, in order to enable the dynamic convolutional layer to adaptively adjust according to the overall information of the sample image, so as to better capture the context and semantic information of the sample image, the weights of the convolutional layer of the dynamic convolutional layer can be determined based on the global feature, and its processing process is shown in formula (7): In formula (7), W dy represents the weights of the convolutional layer of the dynamic convolutional layer; F global represents the first global feature; W a represents the weight matrix, which is used for dot product operation with the first global feature; Softmax () represents a function (such as an activation layer) that converts the input vector into a probability distribution, and is used to generate the weights of the convolutional layer.
[0083] S504, input the second spatial domain feature into the dynamic convolutional layer with the weights of the convolutional layer to obtain the fine-grained feature.
[0084] That is, input the second spatial domain feature into the dynamic convolutional layer with the above-generated weights of the convolutional layer, perform convolution operation, and finally obtain the fine-grained feature.
[0085] In specific implementation, the processing process of obtaining the fine-grained feature can be described by formula (8), as follows: In formula (8), Fdetail represents the fine-grained feature of the sample image; represents the complement of the calculated mask feature, which is used to mask the low-frequency components and extract the high-frequency components; is used to represent extracting the high-frequency components that are not covered by the previous k frequency channels, that is, extracting the high-frequency components in ; represents the element-by-element multiplication of two matrices; DyConv () represents the Dynamic Convolution operation, which is an adaptive convolution operation and can dynamically adjust the weights of the convolutional layer according to the input features. Other parameters have been described above and will not be elaborated here.
[0086] In the disclosed embodiment, the high-frequency component mainly includes the detail information in the sample image. By extracting the high-frequency component, the detail information closely related to the wire fault detection can be extracted, providing more targeted data for subsequent processing. By determining the convolution layer weight of the dynamic convolution layer based on the global information of the first global feature, the convolution operation can be made more targeted and adaptive, so that the dynamic convolution layer can flexibly adjust the convolution operation according to the global situation of the sample image, improve the accuracy of the image detail feature extraction, and further improve the detection accuracy of the wire fault detection model.
[0087] In the embodiment of the present disclosure, in order to integrate feature information at different levels to obtain more comprehensive and richer fusion features, the first global feature, the intermediate semantic feature and the fine-grained feature can be input into the fuser of the model to be trained to obtain the fusion feature, such as Figure 6 As shown, it can be implemented as follows: S601, inputting the first global feature, the intermediate semantic feature and the fine-grained feature into a weight coefficient generator to obtain the weight coefficients corresponding to the first global feature, the intermediate semantic feature and the fine-grained feature respectively.
[0088] In the training phase, the aggregator includes a weight coefficient generator, which is built based on the attention mechanism and consists of a second global average pooling layer, a third global average pooling layer, a multilayer perceptron, and a third activation layer. The first global feature, the intermediate semantic feature, and the fine-grained feature are input into the weight coefficient generator. The weight coefficient generator can calculate the importance weight of each feature based on the attention mechanism, that is, by paying attention to the contribution of each feature to the final classification result, a higher weight is given to the feature with a large contribution, and vice versa, a lower weight is given, so as to obtain the corresponding weight coefficient.
[0089] It should be noted that during the training phase, the weight coefficient generator can dynamically generate weight coefficients based on the input features at different levels, so that the model to be trained can automatically learn how to fuse features at different levels to optimize the objective function. After training, the model to be trained has learned the optimal weights when fusing features at different levels. Therefore, the weight coefficient generator branch can be removed after the training is completed, so that the model to be trained can be simplified and optimized while ensuring the performance of the model, so as to meet the efficiency requirements of the subsequent reasoning stage.
[0090] During implementation, the first global feature, the intermediate semantic feature and the fine-grained feature are input into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature and the fine-grained feature, respectively, which can be implemented based on the following steps: Step B1, using a second global average pooling layer to process the intermediate semantic features to obtain a second global feature; Input the intermediate semantic features into the second global average pooling layer, which performs global average pooling on the intermediate semantic features to obtain the second global feature. The second global feature represents the global overview of the intermediate semantic features and helps to grasp the intermediate semantic features from the overall level when comprehensively considering features at different levels later.
[0091] Step B2: Process the fine-grained features using the third global average pooling layer to obtain the third global feature. Input the fine-grained features into the third global average pooling layer, which performs global average pooling on the fine-grained features to obtain the third global feature. The third global feature reflects the characteristics of the fine-grained features as a whole.
[0092] Step B3: Input the first global feature, the second global feature, and the third global feature into a multi-layer perceptron to obtain multi-layer perceptron features. Input the first global feature, the second global feature, and the third global feature into the multi-layer perceptron. The multi-layer perceptron consists of multiple neural network layers and can perform non-linear transformation and fusion on the input features, and finally output multi-layer perceptron features. The multi-layer perceptron features integrate the information of the three-level features and can further reflect the internal connections between different features.
[0093] Step B4: Input the multi-layer perceptron features into the third activation layer to obtain weight coefficients.
[0094] Input the multi-layer perceptron features into the third activation layer, which performs non-linear transformation through an activation function to obtain weight coefficients.
[0095] In specific implementation, the processing process of obtaining the weight coefficients can be described by formula (9) as follows: In formula (9), are the weight coefficients corresponding to the first global feature, the intermediate semantic features, and the fine-grained features respectively; MLP represents a multi-layer perceptron (i.e., the multi-layer perceptron) used to map the input features to the output space. Softmax The () function is used to convert the input vector into a probability distribution, that is, the weight coefficients corresponding to the first global feature, the intermediate semantic features, and the fine-grained features in the implementation of the present disclosure. The range of the output value is between [0, 1] and the sum is 1. represents performing global average pooling operation on the input; other parameters have been described above and will not be elaborated here.
[0096] In the embodiments of the present disclosure, the first global feature, the intermediate semantic feature, and the fine-grained feature are input into a weight coefficient generator. The weight coefficient generator can, based on the attention mechanism, learn the complex relationships between global features at different levels, enabling the model to be trained to reasonably allocate weight coefficients to these different levels of features when facing different spinning scenarios and types of wire faults that may require different combinations of features for detection. This makes the model to be trained more flexible when fusing different levels of features, thereby improving the accuracy of the finally obtained wire fault detection model in detecting wire faults.
[0097] S602, based on the fuser, perform weighted summation on the first global feature, the intermediate semantic feature, and the fine-grained feature to obtain a fused feature.
[0098] That is, multiply the first global feature, the intermediate semantic feature, and the fine-grained feature by their corresponding weight coefficients and then add them up to obtain the fused feature.
[0099] In specific implementation, when obtaining the fine-grained feature, the implementation process of the fuser can be described by formula (10), as follows: In formula (10), Ffused represents the fused feature; up () is an upsampling operation for adjusting the size of the first global feature to the same size as the intermediate semantic feature and the fine-grained feature; respectively represent the weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature.
[0100] In the embodiments of the present disclosure, the first global feature, the intermediate semantic feature, and the fine-grained feature are input into a weight coefficient generator. The weight coefficient generator can adaptively learn and adjust the weights of each feature according to the input features, enabling the model to be trained to more effectively utilize these features in different situations. Performing weighted summation on the first global feature, the intermediate semantic feature, and the fine-grained feature can integrate feature information at different scales and levels into a unified fused feature, which can describe the features of the sample image from different perspectives and avoid the limitations of a single feature, thereby improving the accuracy of wire fault detection.
[0101] In the embodiments of the present disclosure, to further improve the accuracy of the wire fault detection model in detecting wire faults, the model parameters of the model to be trained can be optimized based on the classification result, such as Figure 7 as shown, and specifically can be implemented as: S701, based on the classification result, determine the detection loss.
[0102] Among them, the detection loss includes at least one of the following: (1)The global loss between the image to be optimized generated based on the first global feature and the sample image; During implementation, a diffusion model can be used to process the first global feature, aiming to recover the sample image based on the first global feature, and an image to be optimized is obtained.
[0103] This global loss focuses on the differences at the overall image level and is used to measure the closeness between the image to be optimized generated by the model and the sample image in terms of global features.
[0104] In the embodiments of the present disclosure, the global loss can be expressed as shown in formula (11): In formula (11), L 1 represents the global loss; ||()||2 represents the L 2 norm (Euclidean norm); F g represents the image to be optimized recovered based on the first global feature; y represents the sample image.
[0105] (2)The intermediate semantic loss; the intermediate semantic loss includes the accumulated value of the Hadamard product of each frequency band mask and the corresponding residual in the mask feature; the residual includes the difference between the intermediate semantic feature and the feature of the target intermediate layer of the teacher model; The intermediate semantic loss is used for knowledge distillation to transfer the knowledge of the intermediate layer features of the teacher model to the intermediate layer features of the current model.
[0106] In the embodiments of the present disclosure, the intermediate semantic loss can be expressed as shown in formula (12): In formula (12), L2 represents the intermediate semantic loss; Mk is the mask feature; Fmid represents the intermediate semantic feature; represents the feature of the target intermediate layer of the teacher model; represents element-wise multiplication (Hadamard product); ||()||1 represents the L1 norm.
[0107] (3)The detail loss; the detail loss is determined based on the total variation regularization term of the fine-grained feature; the total variation regularization term is used to suppress the high-frequency noise in the fine-grained feature; That is, the detail loss pays more attention to the detail level of the sample image, is used to reduce the noise in the fine-grained feature, and at the same time retains the edge information.
[0108] In the embodiments of the present disclosure, the intermediate semantic loss can be expressed as shown in formula (13): In formula (13), L 3 represents the detail loss; TV () represents the total variation regularization term, which is used to reduce noise and incoherence in the detail features; F detail represents the fine-grained feature.
[0109] (4) Classification loss; the classification loss represents the loss value between the classification result and the target classification result.
[0110] S702. Optimize the model parameters of the model to be trained based on the detection loss.
[0111] During implementation, the calculated detection loss can be used to update the parameters of the model to be trained through the backpropagation algorithm, reduce the loss value, and improve the performance of the model to be trained.
[0112] In the embodiments of the present disclosure, the detection loss measures the difference between the model output and the target from different aspects. By comprehensively considering the above-mentioned various detection losses, the model can be constrained and optimized from multiple dimensions such as the global information, local semantics, detail features, and final classification result of the sample image, thereby improving the accuracy of the model to be trained in the silk thread fault detection task.
[0113] In summary, the overall process of the silk thread fault detection model is as Figure 8 shown, including: S801. Perform discrete cosine transform on the sample image to obtain the frequency domain representation of the sample image, where the frequency domain representation includes a zero-frequency component, a low-frequency component, and a high-frequency component.
[0114] S802. Input the sample image into the first global average pooling layer in the zero-frequency enhancer to obtain the global average pooling feature of the sample image; then concatenate the zero-frequency component extracted from the frequency domain representation and the global average pooling feature in the channel dimension to obtain a concatenated feature; finally, input the obtained concatenated feature into the convolutional layer in the zero-frequency enhancer for convolutional operation to obtain the first global feature.
[0115] S803. Input the frequency domain representation into the first activation layer of the learnable frequency domain selector in the low-frequency reconstructor to obtain an activation feature; then input the obtained activation feature into the second activation layer of the learnable frequency domain selector to obtain a frequency weight matrix; then perform inverse discrete cosine transform on the frequency weight matrix to convert the frequency weight matrix from the frequency domain back to the spatial domain to obtain the first spatial domain feature; then multiply the first spatial domain feature and the mask feature element-wise to obtain an intermediate feature; input the intermediate feature into the normalization layer to obtain the intermediate semantic feature.
[0116] S804. Extract the high-frequency components from the frequency-domain representation, convert the high-frequency components to the spatial domain to obtain the second spatial-domain feature, and input the second spatial-domain feature into the dynamic convolution layer in the high-frequency refiner to obtain the fine-grained feature.
[0117] S805. Based on the fuser in the model to be trained, perform weighted summation on the first global feature, intermediate semantic feature, and fine-grained feature to obtain the fused feature.
[0118] S806. Input the fused feature into the classifier to obtain the classification result for the sample image, where the classification result includes the floating filament discrimination result and the broken filament discrimination result.
[0119] S807. Based on the classification result, determine the detection losses such as the global loss, intermediate semantic loss, detail loss, and classification loss between the image to be optimized generated based on the first global feature and the sample image, and optimize the model parameters to be trained based on the detection losses.
[0120] Based on the same concept, an embodiment of the present disclosure proposes a method for detecting filament faults in a spinning process, which is implemented based on the filament fault detection model in the spinning process described above. Specifically, as Figure 9 shown, it can be implemented as: S901. Collect an image of the spinning box to obtain the image to be processed.
[0121] During implementation, an appropriate image acquisition device, such as a high-speed industrial camera, can be used to capture an image of the filament at a specific position of the spinning box as the input data for subsequent filament fault detection models.
[0122] S902. Input the image to be processed into the filament fault detection model obtained through the training method of the filament fault detection model in the spinning process to detect whether there are floating filaments and broken filaments in the filament.
[0123] Input the collected image to be processed into the filament fault detection model obtained through the training method of the filament fault detection model in the spinning process described above. This filament fault detection model performs multi-level feature extraction on the input image, including global features, intermediate semantic features, and fine-grained features, etc. Then, through methods such as the gated fusion mechanism, different-level features are fused to generate a fused feature. Finally, the classifier of the model classifies the fused feature to determine whether there are floating filaments and broken filaments in the image to be processed, and outputs the classification result according to the judgment result.
[0124] The internal processing process of the filament fault detection model is the same as that described above and will not be elaborated here.
[0125] In the embodiments of the present disclosure, by providing a method for detecting silk thread faults in the spinning process, automatic and precise detection of silk thread faults in the spinning process can be achieved, which helps to improve production efficiency, reduce production costs, and improve product quality at the same time.
[0126] Based on the same technical concept, in the embodiments of the present disclosure, a training device 1000 for a silk thread fault detection model in the spinning process is proposed, as Figure 10 shown, including: A frequency domain decomposition module 1001, configured to perform discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; the sample image is obtained by collecting an image of a spinning box; A zero-frequency enhancement module 1002, configured to input the zero-frequency component in the frequency domain representation and the sample image into a zero-frequency enhancer of a model to be trained, and obtain a first global feature of the sample image; A low-frequency reconstruction module 1003, configured to input the frequency domain representation into a low-frequency reconstructor of the model to be trained, and obtain an intermediate semantic feature of the sample image; A high-frequency refinement module 1004, configured to input the frequency domain representation and the first global feature into a high-frequency refiner of the model to be trained, and obtain a fine-grained feature of the sample image; A fusion module 1005, configured to input the first global feature, the intermediate semantic feature, and the fine-grained feature into a fusion unit of the model to be trained, and obtain a fusion feature; A classification module 1006, configured to input the fusion feature into a classifier of the model to be trained, and obtain a classification result for the sample image, where the classification result includes a floating silk discrimination result and a broken silk discrimination result; An optimization module 1007, configured to optimize the model parameters of the model to be trained based on the classification result, and obtain a silk thread fault detection model.
[0127] In some embodiments, the zero-frequency enhancer includes a first global average pooling layer and a convolutional layer; The zero-frequency enhancement module includes: A pooling unit, configured to input the sample image into the first global average pooling layer to obtain a global average pooling feature of the sample image; A splicing unit, configured to splice the zero-frequency component and the global average pooling feature to obtain a spliced feature; A global extraction unit, configured to input the spliced feature into the convolutional layer to obtain a first global feature.
[0128] In some embodiments, the low-frequency reconstructor includes a learnable frequency domain selector and a normalization layer; The low-frequency reconstruction module includes: A weight learning unit, configured to learn the association pattern between different frequency bands in the frequency domain representation based on a learnable frequency domain selector, and generate a frequency weight matrix for frequency selection based on the association pattern; An inverse transformation unit, configured to perform an inverse discrete cosine transform on the frequency weight matrix to obtain a first spatial domain feature; An intermediate feature determination unit, configured to determine an intermediate feature based on the first spatial domain feature and a mask feature; the mask feature includes a band mask corresponding to each frequency in the frequency domain representation; A semantic feature determination unit, configured to input the intermediate feature into a normalization layer to obtain an intermediate semantic feature.
[0129] In some embodiments, the weight learning unit is specifically configured to: Input the frequency domain representation into a first activation layer of the learnable frequency domain selector to obtain an activation feature; the activation feature is used to represent the association pattern between different frequency bands in the frequency domain representation; Input the activation feature into a second activation layer of the learnable frequency domain selector to obtain a frequency weight matrix.
[0130] In some embodiments, the high-frequency refiner includes a dynamic convolution layer; A high-frequency refinement module, including: A high-frequency extraction unit, configured to extract high-frequency components from the frequency domain representation; A conversion unit, configured to convert the high-frequency components to the spatial domain to obtain a second spatial domain feature; A weight determination unit, configured to determine the convolution layer weight of the dynamic convolution layer based on a first global feature; A fine feature determination unit, configured to input the second spatial domain feature into the dynamic convolution layer using the convolution layer weight to obtain a fine-grained feature.
[0131] In some embodiments, the fusion module includes: A weight generation unit, configured to input the first global feature, the intermediate semantic feature, and the fine-grained feature into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature respectively; the weight coefficient generator is constructed based on an attention mechanism; A fusion unit, configured to perform weighted summation on the first global feature, the intermediate semantic feature, and the fine-grained feature based on a fuser to obtain a fusion feature.
[0132] In some embodiments, the weight coefficient generator includes a second global average pooling layer, a third global average pooling layer, a multi-layer perceptron, and a third activation layer; The weight generation unit is specifically configured to: Process the intermediate semantic feature using the second global average pooling layer to obtain a second global feature; The third global average pooling layer is used to process the fine-grained features to obtain third global features; The first global feature, the second global feature, and the third global feature are input into a multi-layer perceptron to obtain multi-layer perceptron features; The multi-layer perceptron features are input into a third activation layer to obtain weight coefficients.
[0133] In some embodiments, the optimization module includes: A loss determination unit for determining a detection loss based on the classification result; An optimization unit for optimizing the model parameters of the model to be trained based on the detection loss; The detection loss includes at least one of the following: The global loss between the image to be optimized generated based on the first global feature and the sample image; The intermediate semantic loss; the intermediate semantic loss includes the accumulated value of the Hadamard product of each frequency band mask and the corresponding residual in the mask feature; the residual includes the difference between the intermediate semantic feature and the feature map of the target intermediate layer of the teacher model; The detail loss; the detail loss is determined based on the total variation regularization term of the fine-grained features; the total variation regularization term is used to suppress high-frequency noise in the fine-grained features; The classification loss; the classification loss represents the loss value between the classification result and the target classification result.
[0134] Based on the same technical concept, an apparatus 1100 for detecting silk thread faults in a spinning process is proposed in an embodiment of the present disclosure, as Figure 11 shown, including: An acquisition module 1101 for acquiring an image of a spinning box to obtain an image to be processed; A detection module 1102 for inputting the image to be processed into a silk thread fault detection model obtained by a training method of a silk thread fault detection model based on a spinning process to detect whether there are problems of floating silk and broken silk in the silk thread.
[0135] For the specific functions and examples of the modules, sub-modules / units of the apparatus in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.
[0136] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0137] Figure 12 It is a structural block diagram of an electronic device according to an embodiment of the present disclosure. As Figure 12As shown, the electronic device includes: a memory 1210 and a processor 1220. The memory 1210 stores a computer program that can run on the processor 1220. The number of the memory 1210 and the processor 1220 can be one or more. The memory 1210 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device executes the method provided in the above method embodiment. The electronic device may further include: a communication interface 1230, configured to communicate with external devices and perform data interaction and transmission.
[0138] If the memory 1210, the processor 1220, and the communication interface 1230 are implemented independently, the memory 1210, the processor 1220, and the communication interface 1230 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 12 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0139] Optionally, in specific implementation, if the memory 1210, the processor 1220, and the communication interface 1230 are integrated on a chip, the memory 1210, the processor 1220, and the communication interface 1230 can communicate with each other through an internal interface.
[0140] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0141] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0142] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (e.g., coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (e.g., infrared, Bluetooth, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., Digital Versatile Disc (DVD)), or a semiconductor medium (e.g., Solid State Disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the present disclosure can be a non-volatile storage medium, in other words, a non-transitory storage medium.
[0143] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0144] In the description of the embodiments of the present disclosure, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0145] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" herein is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0146] In the description of the embodiments of the present disclosure, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0147] The above are only exemplary embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A training method for a silk thread fault detection model in a spinning process, characterized in that, Including: Performing discrete cosine transform on a sample image to obtain a frequency-domain representation of the sample image; The sample image is obtained by collecting an image of a spinning box; Inputting the zero-frequency component in the frequency-domain representation and the sample image into a zero-frequency enhancer of a model to be trained to obtain a first global feature of the sample image; Inputting the frequency-domain representation into a low-frequency reconstructor of the model to be trained to obtain intermediate semantic features of the sample image; Inputting the frequency-domain representation and the first global feature into a high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image; Inputting the first global feature, the intermediate semantic features, and the fine-grained features into a fuser of the model to be trained to obtain a fused feature; Inputting the fused feature into a classifier of the model to be trained to obtain a classification result for the sample image, where the classification result includes a floating filament discrimination result and a broken filament discrimination result; Optimizing model parameters of the model to be trained based on the classification result to obtain a silk thread fault detection model.
2. The method according to claim 1, wherein The zero-frequency enhancer includes a first global average pooling layer and a convolutional layer; The step of inputting the zero-frequency component in the frequency-domain representation and the sample image into a zero-frequency enhancer of a model to be trained to obtain a first global feature of the sample image includes: Inputting the sample image into the first global average pooling layer to obtain a global average pooling feature of the sample image; Concatenating the zero-frequency component and the global average pooling feature to obtain a concatenated feature; Inputting the concatenated feature into the convolutional layer to obtain the first global feature.
3. The method according to claim 1, characterized in that, The low-frequency reconstructor includes a learnable frequency domain selector and a normalization layer; The step of inputting the frequency-domain representation into a low-frequency reconstructor of the model to be trained to obtain intermediate semantic features of the sample image includes: Learning an association pattern between different frequency bands in the frequency-domain representation based on the learnable frequency domain selector, and generating a frequency weight matrix for frequency selection based on the association pattern; Performing inverse discrete cosine transform on the frequency weight matrix to obtain a first spatial domain feature; Determining an intermediate feature based on the first spatial domain feature and a mask feature; the mask feature includes frequency band masks corresponding to each frequency in the frequency-domain representation; Inputting the intermediate feature into the normalization layer to obtain the intermediate semantic features.
4. The method according to claim 3, wherein The step of learning an association pattern between different frequency bands in the frequency-domain representation based on the learnable frequency domain selector, and generating a frequency weight matrix for frequency selection based on the association pattern includes: Inputting the frequency-domain representation into a first activation layer of the learnable frequency domain selector to obtain an activation feature; the activation feature is used to represent the association pattern between different frequency bands in the frequency-domain representation; Inputting the activation feature into a second activation layer of the learnable frequency domain selector to obtain the frequency weight matrix.
5. The method according to any one of claims 1-4, characterized in that, The high-frequency refiner includes a dynamic convolutional layer; The step of inputting the frequency-domain representation and the first global feature into a high-frequency refiner of the model to be trained to obtain fine-grained features of the sample image includes: Extracting high-frequency components from the frequency-domain representation; Convert the high-frequency components to the spatial domain to obtain second spatial domain features; Determine the convolutional layer weights of the dynamic convolutional layer based on the first global feature; Input the second spatial domain features into the dynamic convolutional layer using the convolutional layer weights to obtain the fine-grained features.
6. The method according to claim 1, wherein The step of inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into the fuser of the model to be trained to obtain a fused feature includes: Input the first global feature, the intermediate semantic feature, and the fine-grained feature into a weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature respectively; the weight coefficient generator is constructed based on an attention mechanism; Perform weighted summation on the first global feature, the intermediate semantic feature, and the fine-grained feature based on the fuser to obtain the fused feature.
7. The method according to claim 6, wherein The weight coefficient generator includes a second global average pooling layer, a third global average pooling layer, a multi-layer perceptron, and a third activation layer; The step of inputting the first global feature, the intermediate semantic feature, and the fine-grained feature into the weight coefficient generator to obtain weight coefficients corresponding to the first global feature, the intermediate semantic feature, and the fine-grained feature respectively includes: Process the intermediate semantic feature using the second global average pooling layer to obtain a second global feature; Process the fine-grained feature using the third global average pooling layer to obtain a third global feature; Input the first global feature, the second global feature, and the third global feature into the multi-layer perceptron to obtain a multi-layer perceptron feature; Input the multi-layer perceptron feature into the third activation layer to obtain the weight coefficients.
8. The method according to claim 1, wherein The step of optimizing the model parameters of the model to be trained based on the classification result includes: Determine a detection loss based on the classification result; Optimize the model parameters of the model to be trained based on the detection loss; The detection loss includes at least one of the following: A global loss between the image to be optimized generated based on the first global feature and the sample image; An intermediate semantic loss; the intermediate semantic loss includes the accumulated value of the Hadamard product of each frequency band mask and the corresponding residual in the mask feature; the residual includes the difference between the intermediate semantic feature and the feature map of the target intermediate layer of the teacher model; A detail loss; the detail loss is determined based on the total variation regularization term of the fine-grained feature; the total variation regularization term is used to suppress high-frequency noise in the fine-grained feature; A classification loss; the classification loss represents the loss value between the classification result and the target classification result.
9. A method for detecting silk thread faults in a spinning process, characterized in that, including: Collect an image of the spinning box body to obtain an image to be processed; Input the image to be processed into the silk thread fault detection model obtained by the method according to any one of claims 1-8 to detect whether there are problems of floating silk and broken silk in the silk thread.
10. A training device for a silk thread fault detection model in a spinning process, characterized in that, including: A frequency domain decomposition module for performing a discrete cosine transform on a sample image to obtain a frequency domain representation of the sample image; The sample image is obtained by collecting an image of the spinning box body; A zero-frequency enhancement module, configured to input the zero-frequency component in the frequency-domain representation and the sample image into a zero-frequency enhancer of the model to be trained, and obtain a first global feature of the sample image; A low-frequency reconstruction module, configured to input the frequency-domain representation into a low-frequency reconstructor of the model to be trained, and obtain an intermediate semantic feature of the sample image; A high-frequency refinement module, configured to input the frequency-domain representation and the first global feature into a high-frequency refiner of the model to be trained, and obtain a fine-grained feature of the sample image; A fusion module, configured to input the first global feature, the intermediate semantic feature, and the fine-grained feature into a fuser of the model to be trained, and obtain a fused feature; A classification module, configured to input the fused feature into a classifier of the model to be trained, and obtain a classification result for the sample image, where the classification result includes a floating wire discrimination result and a broken wire discrimination result; An optimization module, configured to optimize model parameters of the model to be trained based on the classification result, and obtain a wire fault detection model.
11. A silk thread fault detection device in a spinning process, characterized in that, including: An acquisition module, configured to perform image acquisition on a spinning box to obtain an image to be processed; A detection module, configured to input the image to be processed into the wire fault detection model obtained by the device according to claim 10, to detect whether there are floating wire and broken wire problems in the wire.
12. An electronic device, characterized in that, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method according to any one of claims 1-9.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
14. A computer program product, characterized in that, including a computer program, which when executed by a processor, implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Image spoofing detection method and device and computer readable storage medium
CN117523217A
Two-dimensional feature selection verification method for optical remote sensing image similarity evaluation
CN118506135A
Remote sensing image segmentation method based on channel enhancement and cross-level multi-input features
CN119380018A
Mechanical rotating part fault diagnosis method, device, equipment and medium
CN119917936A
Methods and apparatuses for video segmentation, classification, and retrieval using image class statistical models
US20020028021A1