Embroidery stitch identification method based on deep learning
By improving the YOLO model and combining the multi-scale texture preprocessing module and the direction perception attention module, the problem of low recognition accuracy in embroidery needle recognition is solved, and high accuracy recognition under complex conditions is achieved.
Patent Information
- Application Number
- CN202510452471.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to fully capture the subtle differences in embroidery texture in embroidery needle recognition, especially when the light changes, needle and thread color differences or the embroidery angle offset, the recognition accuracy significantly decreases.
Using a deep learning-based method, the feature extraction and classification of embroidered pictures is performed by improving the YOLO model, combining multi-scale texture preprocessing module, texture perception cross-stage part module and direction perception attention module.
It significantly improves the accuracy and robustness of embroidery needle recognition, and can maintain a high recognition accuracy under complex textures, lighting changes and diverse embroidery angles.
Smart Images

Figure CN120014368A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular to an embroidery needle recognition method based on deep learning. Background Art
[0002] Automatic recognition technology for embroidery stitches is the core link of the digitization and intelligence of embroidery technology. It has important application value in the fields of cultural heritage protection, intelligent embroidery equipment control, product quality inspection, etc. Traditional methods mainly rely on manual feature extraction (such as grayscale co-occurrence matrix, Fourier transform, etc.) to achieve stitch classification through texture analysis. However, embroidery textures have complex three-dimensional structures, tiny details and irregular stitch distribution. Methods based on manual feature extraction (such as grayscale co-occurrence matrix GLCM, Fourier transform, etc.) achieve stitch classification by designing specific texture features (such as energy, moment of inertia). Although this type of method has a fast calculation speed, it is difficult to fully capture the subtle differences in embroidery textures because the feature design relies on manual experience. In particular, when the lighting changes, the needle and thread color is different, or the embroidery angle is offset, the recognition accuracy drops significantly.
[0003] Although the gray-level co-occurrence matrix (GLCM) is widely used in extracting texture features, it has its limitations. For example, Haralick et al. (Reference: Haralick, RM, et al. "Textural features for image classification." IEEE Transactions on Systems, Man, and Cybernetics 3.6(1973): 610-621.) first proposed the concept of GLCM in 1973 to extract texture features of images. However, the computational complexity of GLCM is high, especially when processing high-resolution images, the calculation takes a long time. In addition, GLCM is sensitive to noise, which can significantly affect the extraction effect of its texture features. The features extracted by GLCM may be redundant, and the ability to capture local texture changes is weak. When processing complex textures, GLCM may not be able to fully describe its texture characteristics, especially when the texture has multi-scale and multi-directional features. In embroidery stitch recognition, the stitch directions are diverse (such as back stitch and rolling stitch), and a single angle parameter is difficult to cover all situations, resulting in limited feature expression capabilities. In addition, GLCM is calculated based on grayscale images. If the contrast between the embroidery thread color and the background is low (such as light-colored thread embroidered on white cloth), the segmentation error will directly affect feature extraction. This leads to high requirements and low fault tolerance for the establishment of the data set when processing embroidery stitch images, and is likely to cause errors.
[0004] The Fourier Transformation (FFT) aims to convert the image from the spatial domain to the frequency domain and extract global texture features through the spectral energy distribution (low frequency corresponds to smooth areas, high frequency corresponds to details). Although it can suppress high-frequency noise, it is suitable for images with uneven light or slight stains through low-pass filtering. However, the Fourier transform converts the image as a whole into a frequency domain that reflects the overall spectral distribution. Therefore, the key differences in embroidery needles (such as the chain stitches of the chain stitch, the granularity of the seed needle, etc.) are often reflected in the local stitch structure, and the texture of the texture is very different from the texture of the texture. FFT has difficulty locating these local features, resulting in confusion between similar stitches. At the same time, the images of embroidery stitches are often rotated during the embroidery process due to changes in the embroidery angle (such as the 45° direction of the oblique needle). And the Fourier transform spectrum is very sensitive to rotation, and the phase information will change significantly even if it is slightly rotated. This requires additional registration steps (such as polar coordinate transformation) to restore directional consistency, which increases the computational burden. In contrast, convolutional neural networks are able to maintain a high degree of robustness, thanks to the translation invariance of the convolution kernel. Summary of the invention
[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide an embroidery needle recognition method based on deep learning in view of the shortcomings of the prior art, comprising the following steps:
[0006] Step 1, determine the needle method and collect relevant pictures. Generally, the needle method is determined according to the needle method to be identified, and then the corresponding needle method pictures are collected online or found in books according to the needle method required;
[0007] Step 2: preprocess and annotate the images to obtain an embroidery dataset;
[0008] Step 3: Divide the embroidery dataset into training set, validation set and test set, and write the configuration;
[0009] Step 4, improving the YOLO model to obtain an improved YOLO model;
[0010] Step 5, use the training set to train the improved YOLO model;
[0011] Step 6: Use the validation set to evaluate the improved YOLO model. If the conditions are not met, such as the precision, recall or mAP value is too low, such as less than 0.3, the conditions are not met, and the model structure and model training parameters are adjusted, and return to step 5. If the conditions are met, input the test set into the improved YOLO model to obtain the prediction results.
[0012] In step 2, the embroidery data set includes embroidery stitch categories of satin embroidery, rolling needle embroidery, flat stitch, grab stitch, loop stitch, herringbone embroidery, seed embroidery, French knot, return, random stitch, embroidery lock, long and short stitch, and straight stitch, and each stitch category contains a corresponding embroidery picture.
[0013] In step 2, the preprocessing includes: performing data enhancement processing on the embroidery image using a rotation mirror method, and superimposing Gaussian noise on the embroidery image: , in, This is the embroidery picture after adding Gaussian noise. For the original embroidery picture, The mean is 0 and the variance is Gaussian noise.
[0014] In step 3, the configuration writing includes: Configure the data set: configure the location of the training set, validation set, and test set, as well as the category labels of the stitches in mydata.yaml; for example, category 0 is seed stitch, and category 2 is straight stitch;
[0015] A configuration file of the model structure is created. The total number of needle method categories is determined in the configuration file, and the network structure of the model is defined (for example, which modules the image should pass through in the backbone module). The model name is model.yaml.
[0016] Step 4 includes: The network structure of the YOLO model includes a backbone network Backbone, a neck network Neck and a detection head Head.
[0017] In step 4, the backbone network is used to extract feature I from the input image, the neck network is used to aggregate and enhance features of different scales, and the detection head classifies and locates the target according to the extracted features.
[0018] In step 4, a multi-scale texture preprocessing module is inserted at the front end of the backbone network Backbone to optimize the input features. The multi-scale texture preprocessing module converts the embroidery image to the HSV (Hue, Saturation, Value) color space through illumination normalization for processing, and then converts it into a red, green, and blue RGB image. In the HSV space, the brightness V channel is localized for contrast normalization, and the local mean and standard deviation are calculated using a sliding window, and α and β are dynamically adjusted to retain the details of the reflective area:
[0019] ,
[0020] in, is the pixel value of the embroidery image at position (x, y), is the mean value of all pixels in the local window Ω centered at position (x, y), is the standard deviation of the pixel values within the local window Ω centered at position (x, y); is a constant; is the scaling factor, which adjusts the dynamic range of the normalized feature and is initially set to 1; is the offset coefficient, which controls the overall brightness of the normalized feature and is initially set to 0;
[0021] Then, the multi-scale texture preprocessing module performs non-photorealistic rendering (NPR) operation, uses a multi-directional Gabor filter to extract the texture direction information in the embroidery image, and strengthens or weakens the edge and texture. The formula is:
[0022] ,
[0023] ,
[0024] ,
[0025] in, The direction is Gabor filter; is the smoothing strength coefficient, usually ranging from 0.5 to 1.0; is the edge enhancement coefficient, usually 0.5~1.0; T is the adaptive edge threshold, usually 100~150; is position-wise multiplication; ; is the position (x, y) in the input image I Direction: Response map obtained by applying multi-directional Gabor filter; represents the sum of the absolute values of the Gabor filter responses in all directions at the position (x, y) in the image. Represents the set of all directions involved in the calculation, Represents the final image obtained after non-photorealistic rendering.
[0026] Step 4 also includes: replacing the shallow C2f module of the backbone network Backbone with the texture-aware cross-stage partial module TextureAwareCSP, and optimizing the data flow of the embroidery feature map X obtained by the multi-scale texture preprocessing module and the convolution operation through dual-branch parallel processing: the input embroidery feature map is first split into a high-frequency texture branch and a low-frequency semantic branch. The high-frequency texture branch extracts the stitch edge details through 3×3 convolution, Hard-Swish activation function and multi-directional Gabor filtering, retains the original resolution, and finally obtains the high-frequency feature map H of the embroidery; the low-frequency semantic branch compresses the resolution through 3×3 convolution downsampling convolution with a step size of 2, and uses Ghost convolution to capture global semantic information, and finally outputs the low-frequency feature map L of the embroidery;
[0027] The high-frequency texture branch is represented as:
[0028] ,
[0029] The X size of the embroidery feature map is , C is the number of channels, H is the height of the feature map, and W is the width of the feature map; yes The convolution of The output channel is C / 2; is the Hard-Swish activation function, It is a multi-directional Gabor convolution group;
[0030] The low-frequency semantic branch is expressed as:
[0031] ,
[0032] Where S is the step size, express Downsampling convolution with stride 2, It is Ghost convolution;
[0033] The embroidery features obtained by the high-frequency texture branch and the low-frequency semantic branch are spliced after upsampling, alignment, and dynamic weighted fusion of high-frequency texture and low-frequency semantic information through the channel attention mechanism, and finally the processed embroidery feature map is output. :
[0034] ,
[0035] in, The low-frequency feature map L of the embroidery is upsampled by bilinear interpolation to Dimension, Concat means connection, is the channel attention mechanism, and the calculation formula is:
[0036] ;
[0037] in It is channel-by-channel multiplication, GAP is global average pooling; F is the feature map of the input channel attention mechanism, which here represents the feature map of the concatenation of the high-frequency feature H and the low-frequency feature L of the embroidery; MLP (Multi-Layer Perceptron) is a multi-layer perceptron;
[0038] Add the OrientationAttention module to the neck network to enhance the sensitivity to the direction of embroidery stitch arrangement;
[0039] The OrientationAttention module performs the following operations: the input embroidery feature map X is firstly convolved by 1×1, and the number of output channels K is output. The input embroidery feature is mapped to the direction weight space to generate a multi-directional weight map, and the attention weights of different directions of each spatial position are obtained by Softmax normalization;
[0040] ,
[0041] Where X∈R C×H×W , R represents the real number space, yes The convolutional layer; is the obtained multi-directional attention weight map;
[0042] Represents normalization along the direction dimension K, ensuring ; Represents the attention strength of the kth direction at position (x, y);
[0043] Perform multi-directional convolution on the input embroidery feature map X to generate direction-sensitive features:
[0044] ,
[0045] in is the feature map of the kth direction, k=1, 2,…, K; is the learnable convolution kernel in the kth direction; ∗ is the standard convolution operation;
[0046] Fusion of multi-directional features by directional weights to enhance dominant directional response:
[0047] ,
[0048] in represents the attention weight of the kth direction,
[0049] Finally, the enhanced directional features are superimposed on the original features through residual connections:
[0050] ,
[0051] in, is the learnable scaling factor; It is the final output embroidery direction enhanced feature map.
[0052] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0053] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.
[0054] The method of the present invention is mainly used in the field of embroidery technology digitization and intelligent textile, especially in the protection of cultural heritage and intelligent embroidery equipment. This technology performs end-to-end feature extraction and classification on embroidery pictures acquired by high-resolution scanners or cameras through deep learning models (such as convolutional neural networks CNN and attention mechanisms), and can accurately identify different needle types (such as flat stitches, grab stitches, seed stitches, etc.). Compared with traditional manual feature extraction methods, deep learning methods can automatically learn the multi-level features of embroidery textures, significantly improve recognition accuracy and robustness, especially in the case of complex textures, lighting changes and embroidery angle diversity. This is of great significance for the digital archiving of embroidery technology, the intelligent generation of embroidery patterns, the quality inspection of embroidery products, and the inheritance and innovation of traditional embroidery skills.
[0055] This technology can be applied in the following fields:
[0056] Digital protection of cultural heritage: Through deep learning models (such as convolutional neural networks), high-precision needle identification of embroidery artifacts can be performed, which can achieve digital archiving and restoration of endangered embroidery needle techniques, provide automated analysis tools for museums and research institutions, and assist in the protection and inheritance of traditional crafts.
[0057] Intelligent embroidery equipment control: In automated embroidery equipment, deep learning-based needle recognition technology can analyze and identify the needle types in the design pattern in real time, thereby automatically generating machine embroidery instructions, significantly improving the embroidery efficiency and accuracy of complex patterns and reducing manual programming costs.
[0058] Textile product quality inspection: Using deep learning models to perform end-to-end analysis of embroidery product images can quickly detect defects such as needle errors and stitch deviations, realize intelligent quality inspection of the production line, reduce manual sampling errors, and ensure high-standard and high-precision output of embroidery products.
[0059] Application in the field of education: In the teaching of embroidery skills, the interactive system driven by deep learning can analyze the needlework execution of embroidery students (embroidery enthusiasts) in real time, provide visual feedback and error correction suggestions, and improve teaching efficiency and learning experience.
[0060] Beneficial effects: Compared with the gray-level co-occurrence matrix (GLCM) and Fourier transform (FFT) methods commonly used in traditional embroidery stitch recognition, the method of the present invention can more accurately identify the stitch category in embroidery pictures. Traditional methods often have problems such as low recognition accuracy, sensitivity to noise, and feature redundancy when dealing with complex textures, lighting changes, and diversity of embroidery angles. The method of the present invention can automatically learn the multi-level characteristics of embroidery textures through a deep learning model, and has significantly improved recognition and toughness. This solution has a wider applicability and can maintain high recognition accuracy and robustness under complex textures, lighting changes, and diversity of embroidery angles. It is suitable for a variety of embroidery stitches (such as flat stitches, grab stitches, seed stitches, etc.), and can accurately and effectively identify embroidery pictures under different colors, backgrounds, and lighting conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a flow chart of the present invention.
[0062] Figure 2 This is an identification picture of rolling needle embroidery.
[0063] Figure 3 This is an identification diagram of backstitch embroidery.
[0064] Figure 4 This is an identification diagram of fishbone embroidery stitch. DETAILED DESCRIPTION
[0065] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0066] like Figure 1 As shown, an embodiment of the present invention provides an embroidery stitch recognition method based on deep learning. The method of the present invention realizes the recognition of embroidery pictures based on the YOLO model.
[0067] The embroidery dataset contains embroidery stitches such as rolling needle embroidery, fishbone embroidery, and backstitch. Each stitch category contains more than 100 different pictures. After completing the collection of basic embroidery pictures, in order to increase the diversity of training data and thus improve the generalization ability of the model, data enhancement processing is required before labeling. In addition to the basic rotation and mirroring method, you can also choose to add noise. Superimposing Gaussian noise on the image can better increase the complexity and uncertainty of the data. It helps the model learn the intrinsic characteristics of the embroidery samples more robustly and avoids overfitting problems caused by clean (noise-free) images. For example, you can choose (relative to the image pixel value range of 0~255), that is, the standard deviation accounts for about 0.4% of the original image mean.
[0068] ,
[0069] in This is the image after adding noise. For the original picture, The mean is 0 and the variance is Gaussian noise is often used to simulate image jitter or degradation, in which Gaussian noise is a special image.
[0070] Then the dataset is labeled. After the labeling is completed, the embroidery dataset is divided into training set, validation set and test set in a ratio of 7:2:1. After the dataset is divided into training set, validation set and test set, configuration writing is required. Configuration writing is divided into two parts. One is the configuration writing of the dataset, which needs to configure the location of the training set, validation set, and test set, as well as the category label of the model, named mydata.yaml; the other is the configuration file of the model structure, in which the total number of stitch categories needs to be determined and the network structure of the model needs to be defined, named model.yaml.
[0071] The network structure of the YOLO model is mainly composed of the backbone network (Backbone), the neck network (Neck) and the detection head (Head). The backbone network is responsible for extracting features from the input image, the neck network is used to aggregate and enhance features of different scales to improve the model's ability to detect multi-scale targets, and the detection head is responsible for classifying and locating targets based on the extracted features;
[0072] Based on the Backbone module of YOLO11, a multi-scale texture preprocessing module is inserted at the front end of the entire Backbone to optimize the input features, so that the Backbone can more efficiently extract the texture information related to the stitches. This module explicitly enhances the microscopic directional features and macroscopic structural consistency of the embroidery image through illumination normalization → texture abstraction.
[0073] Lighting normalization converts the embroidery image to HSV space for processing, and then converts it to RGB image. In HSV space, the brightness (V channel) is normalized for local contrast. The local mean and standard deviation are calculated using a sliding window and α and β are dynamically adjusted to retain the details of the reflective area:
[0074] ,
[0075] in, is the pixel value of the original embroidery image at position (x, y), The mean of all pixels within the local window Ω centered at pixel (x, y), is the standard deviation of the pixel values within the local window Ω centered at pixel (x, y); is a small constant that prevents division by 0, such as 10 -6 .
[0076] Then, a non-realistic rendering (NPR) operation is performed to extract the texture direction information in the embroidery image using a multi-directional Gabor filter, and to strengthen or weaken the edges / textures in a specific direction, thereby highlighting the details of the embroidery needlework, similar to the effect of sketching style or edge enhancement. Exponential smoothing is performed on non-edge areas to retain low-frequency textures and linear enhancement is performed on significant edges.
[0077] ,
[0078] ,
[0079] ,
[0080] in, is a Gabor filter with orientation ( ), is the smoothing strength coefficient, usually ranging from 0.5 to 1.0. is the edge enhancement coefficient, usually ranging from 20 to 50, T is the adaptive edge threshold, .
[0081] The C2f module in the shallow layer (high-resolution stage) of the backbone is replaced with TextureAwareCSP to enhance the ability to capture detailed textures. The embroidery feature map optimizes the data flow through dual-branch parallel processing: the input features are first split into a high-frequency texture branch and a low-frequency semantic branch. The high-frequency texture branch extracts stitch edge details through 3×3 convolution and multi-directional Gabor filtering, retaining the original resolution; the low-frequency semantic branch compresses the resolution through downsampling convolution with a step size of 2, and uses lightweight Ghost convolution to capture global semantics. The two-way features are spliced after upsampling alignment, and the high-frequency texture and low-frequency semantic information are dynamically weighted and fused through the channel attention mechanism, and finally optimized and output by the C2f module.
[0082] High frequency texture branch:
[0083] ,
[0084] Where X is the input embroidery feature map, the size is , yes The convolution output channel is C / 2; is the Hard-Swish activation function, is a multi-directional Gabor convolution group,
[0085] Low-frequency semantic branch:
[0086] ,
[0087] Among them, S=2 is the downsampling convolution with a step size of 2, It is Ghost convolution, which is used to reduce the amount of calculation;
[0088] Feature Fusion:
[0089] ,
[0090] in, is bilinear interpolation upsampling to , is the channel attention mechanism, and the calculation formula is:
[0091] ;
[0092] in It is channel-by-channel multiplication, and GAP is global average pooling.
[0093] In order to meet the detection requirements of dense small targets and complex texture features in the embroidery stitch recognition task, an OrientationAttention module is added to the Neck part to enhance the sensitivity to the arrangement direction of embroidery stitches.
[0094] The OrientationAttention module explicitly models the directional distribution of features through directional weight generation and dynamic directional convolution.
[0095] In the embroidery data processing flow of the OrientationAttention module, the input embroidery feature map X first passes through a 1×1 convolution, and the output channel number K is used to map the input embroidery feature to the direction weight space, generate a multi-directional weight map (the direction cardinality K is predefined, such as 8 directions), and obtain the attention weights of different directions of each spatial position through Softmax normalization;
[0096] ,
[0097] Where X∈R C×H×W , yes The convolution layer has an output channel of K, where K is the direction cardinality;
[0098] is normalized along the direction dimension (K), ensuring ;
[0099] Perform multi-directional convolution on the input embroidery feature X to generate direction-sensitive features.
[0100] ,
[0101] Where k = 1, 2, ..., K; is the learnable convolution kernel in the kth direction; ∗ is the standard convolution operation, using the same padding to maintain the resolution.
[0102] Multi-directional features are fused according to directional weights to enhance the dominant direction response.
[0103] ,
[0104] in is a position-wise multiplication, represents the attention weight of the kth direction (from middle slice), is the feature map in the kth direction.
[0105] Finally, through the residual connection (including learnable scaling coefficients ) The enhanced directional features are superimposed on the original features, and the output maintains the same resolution and number of channels as the input, retaining the original features and preventing directional enhancement from destroying the existing semantics.
[0106] ,
[0107] in, is a learnable scaling factor (initialized to 0.1) that controls the strength of directional enhancement.
[0108] Result evaluation: To identify acupuncture techniques from images, it is necessary not only to correctly predict the category of the acupuncture techniques, but also to predict the specific location of the acupuncture techniques. Result evaluation can be divided into two parts: one is the prediction index of the prediction box, and the other is the classification prediction index.
[0109] 1. Evaluation indicator of the prediction box: intersection over union (IoU)
[0110] The accuracy of the prediction box is reflected by IOU. IOU is an important indicator in the object detection problem. It reflects the degree of overlap between the labeled box and the prediction box during the training phase and is used to measure the accuracy of the prediction box.
[0111] ,
[0112] Among them, Y1 represents the union area of the predicted box and the true box, and Y2 represents the intersection area of the predicted box and the true box. The higher the IoU value, the greater the overlap between the predicted box and the true box, and the higher the regional accuracy of the needle method. Usually, the IoU threshold is set to 0.5, that is, when IoU ≥ 0.5, the prediction is considered to be a positive example.
[0113] 2. Classification prediction indicators
[0114] Confusion Matrix: It is used to show the relationship between the model prediction results and the true labels. It is usually divided into four categories: True Positive (TP), False Positive (FP), True Negative (TN), and False Negative (FN). TP indicates the number of samples where the actual needle method is consistent with the predicted needle method. FP indicates that the model classifies the background or other needle method categories into the current needle method category. FN indicates that the model failed to correctly detect the real needle method category, resulting in the target not being recognized or misclassified into other categories.
[0115] Through the confusion matrix, the following indicators can be calculated: Precision, Recall, F1-score, PR curve,
[0116] Precision: Among all the samples predicted as a certain stitch category, it is the proportion of the actual correct stitch category. The calculation formula is:
[0117] ,
[0118] Recall: Among all the samples that are actually a certain stitch category, it is the proportion that is correctly predicted as the correct stitch category. The calculation formula is:
[0119] ,
[0120] F1-Score: The harmonic mean of precision and recall, used to comprehensively measure the classification performance of the model. The calculation formula is:
[0121] ,
[0122] With the prediction metrics of the prediction box and the classification prediction metrics, the two can be combined into the evaluation metrics of the model:
[0123] mAP@0.5: Calculate the average precision of each category when the IoU threshold is 0.5.
[0124] mAP@[0.5:0.95]: Calculate the average precision of each category in the range of IoU thresholds from 0.5 to 0.95 and take the average value.
[0125] The higher the mAP value, the better the comprehensive detection performance of the model.
[0126] Metrics such as IoU, confusion matrix, precision, recall, F1-score, and mAP together constitute a comprehensive evaluation system for the performance of the object detection model. In practical applications, these metrics should be comprehensively considered according to the requirements of specific tasks to comprehensively evaluate the detection ability of the model.
[0127] The implementation results of the method of the present invention are as shown in Figure 2 , Figure 3 , Figure 4 ( Figure 2 The English in it is the pinyin "rolling needle embroidery", Figure 3 The English in it is the pinyin "back stitch", Figure 4 The English in it is the pinyin "fishbone stitch"), Figure 2 , Figure 3 , Figure 4 show the recognition diagrams of different stitch embroideries. The numbers in the diagrams represent the confidence levels of the corresponding stitches. This reflects the end-to-end feature extraction of embroidery pictures, which can identify different stitch categories, providing an efficient, robust, and interpretable solution for the digitalization and automated quality inspection of embroidery processes.
[0128] The present invention provides an embroidery needle recognition method based on deep learning. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A method for embroidery needle recognition based on deep learning, characterized in that: The following steps are involved: Step 1, determine the needle method and collect relevant pictures. Generally, the needle method is determined according to the needle method to be identified, and then the corresponding needle method pictures are collected online or found in books according to the needle method required; Step 2: preprocess and annotate the images to obtain an embroidery dataset; Step 3: Divide the embroidery dataset into training set, validation set and test set, and write the configuration; Step 4, improving the YOLO model to obtain an improved YOLO model; Step 5, use the training set to train the improved YOLO model; Step 6: Use the validation set to evaluate the improved YOLO model. If the conditions are not met, adjust the model structure and model training parameters and return to step 5. If the conditions are met, input the test set into the improved YOLO model to obtain the prediction results.
2. The method according to claim 1, characterized in that In step 2, the embroidery data set includes embroidery stitch categories of satin embroidery, rolling needle embroidery, flat stitch, grab stitch, loop stitch, herringbone embroidery, seed embroidery, French knot, return, random stitch, embroidery lock, long and short stitch, and straight stitch, and each stitch category contains a corresponding embroidery picture.
3. The method according to claim 2, characterized in that In step 2, the preprocessing includes: performing data enhancement processing on the embroidery image using a rotation mirror method, and superimposing Gaussian noise on the embroidery image: , in, This is the embroidery picture after adding Gaussian noise. For the original embroidery picture, The mean is 0 and the variance is Gaussian noise.
4. The method according to claim 3, characterized in that In step 3, the configuration writing includes: Write the configuration of the data set: configure the location of the training set, validation set, and test set, as well as the category label mydata.yaml of the needle method; A configuration file of the model structure is established, in which the total number of needle method categories is determined and the network structure of the model is defined. The model name is model.yaml.
5. The method according to claim 4, characterized in that Step 4 includes: The network structure of the YOLO model includes a backbone network Backbone, a neck network Neck and a detection head Head.
6. The method according to claim 5, characterized in that In step 4, the backbone network is used to extract feature I from the input image, the neck network is used to aggregate and enhance features of different scales, and the detection head classifies and locates the target according to the extracted features.
7. The method according to claim 6, characterized in that In step 4, a multi-scale texture preprocessing module is inserted at the front end of the backbone network Backbone to optimize the input features. The multi-scale texture preprocessing module converts the embroidery image to the HSV color space through illumination normalization for processing, and then converts it into a red, green and blue RGB image after processing. In the HSV space, the brightness V channel is localized for contrast normalization, the local mean and standard deviation are calculated using a sliding window, and α and β are dynamically adjusted to retain the details of the reflective area: , in, is the pixel value of the embroidery image at position (x, y), is the mean value of all pixels in the local window Ω centered at position (x, y), is the standard deviation of the pixel values within the local window Ω centered at position (x, y); is a constant; is the scaling factor; is the offset coefficient; Then, the multi-scale texture preprocessing module performs non-realistic rendering operations, uses a multi-directional Gabor filter to extract the texture direction information in the embroidery image, and strengthens or weakens the edge and texture. The formula is: , , , in, The direction is Gabor filter; is the smoothing strength coefficient; is the edge enhancement coefficient; T is the adaptive edge threshold; is position-wise multiplication; ; is the position (x, y) in the input image I Direction: Response map obtained by applying multi-directional Gabor filter; represents the sum of the absolute values of the Gabor filter responses in all directions at the position (x, y) in the image. Represents the set of all directions involved in the calculation, Represents the final image obtained after non-photorealistic rendering.
8. The method according to claim 7, characterized in that Step 4 also includes: replacing the shallow C2f module of the backbone network Backbone with the texture-aware cross-stage partial module TextureAwareCSP, and optimizing the data flow of the embroidery feature map X obtained by the multi-scale texture preprocessing module and the convolution operation through dual-branch parallel processing: the input embroidery feature map is first split into a high-frequency texture branch and a low-frequency semantic branch. The high-frequency texture branch extracts the stitch edge details through 3×3 convolution, Hard-Swish activation function and multi-directional Gabor filtering, retains the original resolution, and finally obtains the high-frequency feature map H of the embroidery; the low-frequency semantic branch compresses the resolution through 3×3 convolution downsampling convolution with a step size of 2, and uses Ghost convolution to capture global semantic information, and finally outputs the low-frequency feature map L of the embroidery; The high-frequency texture branch is represented as: , The X size of the embroidery feature map is , C is the number of channels, H is the height of the feature map, and W is the width of the feature map; yes The convolution of The output channel is C / 2; is the Hard-Swish activation function, It is a multi-directional Gabor convolution group; The low-frequency semantic branch is expressed as: , Where S is the step size, express Downsampling convolution with stride 2, It is Ghost convolution; The embroidery features obtained by the high-frequency texture branch and the low-frequency semantic branch are spliced after upsampling, alignment, and dynamic weighted fusion of high-frequency texture and low-frequency semantic information through the channel attention mechanism, and finally the processed embroidery feature map is output. : , in, The low-frequency feature map L of the embroidery is upsampled by bilinear interpolation to Dimension, Concat means connection, is the channel attention mechanism, and the calculation formula is: ; in is channel-by-channel multiplication, GAP is global average pooling; F is the feature map of the input channel attention mechanism; MLP is a multi-layer perceptron; Add the OrientationAttention module to the neck network to enhance the sensitivity to the direction of embroidery stitch arrangement; The OrientationAttention module performs the following operations: the input embroidery feature map X is firstly convolved by 1×1, and the number of output channels K is output. The input embroidery feature is mapped to the direction weight space to generate a multi-directional weight map, and the attention weights of different directions of each spatial position are obtained by Softmax normalization; , Where X∈R C×H×W , R represents the real number space, yes The convolutional layer; is the obtained multi-directional attention weight map; Represents normalization along the direction dimension K, ensuring ; Represents the attention strength of the kth direction at position (x, y); Perform multi-directional convolution on the input embroidery feature map X to generate direction-sensitive features: , in is the feature map of the kth direction, k=1, 2,…, K; is the learnable convolution kernel in the kth direction; ∗ is the standard convolution operation; Fusion of multi-directional features by directional weights to enhance dominant directional response: , in represents the attention weight of the kth direction, Finally, the enhanced directional features are superimposed on the original features through residual connections: , in, is the learnable scaling factor; It is the final output embroidery direction enhanced feature map.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Cited By
Forward-looking sonar image obstacle detection method, device and equipment and storage medium
CN120726464A
A method, apparatus, device, and storage medium for obstacle detection using forward-looking sonar images.
CN120726464B