Light-weight real-time dressing detection method for advanced manufacturing industry

By improving the YOLOv10 model and introducing a lightweight network and attention mechanism, the shortcomings in applicability, speed and accuracy of the dress detection algorithm in the existing technology are solved, and efficient and real-time dress detection effects are achieved in an advanced manufacturing environment.

CN119963487APending Publication Date: 2025-05-09TIANJIN IRISTAR TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411963270.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing dress compliance detection algorithms have shortcomings in applicability, speed and accuracy, especially when facing different dress code differences in different industries, computing power limitations under cost control, and small target detection difficulties.

Method used

A lightweight real-time dress detection method for advanced manufacturing is proposed. By collecting and preprocessing data, the YOLOv10 model is improved, MobileNetV4, dynamic convolution and attention mechanism is introduced, and small object detection heads and DA attention mechanisms are added to improve detection accuracy and speed.

Benefits of technology

While maintaining detection accuracy, it significantly improves the real-time performance of the algorithm and the accuracy of small target detection, and meets the real-time detection needs of the advanced manufacturing production environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963487A_ABST
    Figure CN119963487A_ABST
Patent Text Reader

Abstract

The invention provides an advanced manufacturing industry-oriented lightweight real-time dressing detection method, which comprises the following steps of: collecting dressing data images, and constructing a dressing image data set; performing data enhancement and preprocessing on the collected data; the YOLOv10 is used as a baseline model, the YOLOv10 is improved, and a dressing detection model is constructed; sending the data into a dressing detection model for training and testing; and carrying out dressing detection by using the trained dressing detection model. The method can better adapt to different types of advanced manufacturing industry production environments; a lightweight backbone network and an attention mechanism are adopted, the detection precision and the real-time performance of the algorithm are improved, the recognition precision and the user experience are greatly improved, and the method has great production practice significance. According to the invention, dressing detection can be carried out on personnel in a production environment in real time. And the small target detection mode is optimized, and the accuracy and robustness of the algorithm on small target detection are improved, so that the algorithm precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer technology and target detection technology, and in particular relates to a lightweight real-time clothing detection method for advanced manufacturing industries. Background Art

[0002] In my country's advanced manufacturing industries such as medicine, semiconductors, food production clean rooms, laboratories, chemical production, etc., hygiene and safety standards are of paramount importance to ensure product quality. These industries usually require workers to strictly abide by the corresponding dress code, such as wearing dust-free hats, masks, protective clothing and gloves, etc., to reduce the risk of contamination by microorganisms and particulate matter. However, in the actual production process, non-compliant dress may pollute the production environment, thereby threatening the safety of the products produced.

[0003] With the development of deep learning technology, the dress inspection method based on manual sampling has been gradually replaced by the dress compliance detection algorithm based on deep learning due to its low efficiency, insufficient accuracy and lack of real-time performance. The existing dress compliance detection algorithms are mainly composed of human body and key point detection algorithms and dress detection algorithms. The dress detection algorithm based on target detection is the core algorithm for detecting and judging whether the target is dressed in compliance. Its speed and accuracy directly affect the overall effect of the dress compliance detection algorithm.

[0004] However, there are several problems in the actual application of clothing detection algorithm:

[0005] 1. There are differences in dress code requirements and types of clothing between different industries, such as medicine, semiconductors, and food. Even for the same type of clothing requirements, such as protective clothing, there are differences in styles, which limits the applicability of algorithm detection.

[0006] 2. In order to control costs, the current dress compliance detection algorithm is mainly deployed on edge computing devices, whose computing power is much lower than that of servers. Therefore, the algorithm reasoning speed is very high, which requires the model parameters to be small and lightweight enough, but the detection accuracy will be reduced, which makes it difficult for the algorithm to have a good balance between speed and accuracy.

[0007] 3. When the objects in the scene are small, such as masks or gloves, the algorithm has difficulty detecting small objects.

[0008] Therefore, when the clothing detection algorithm is actually used, due to the wide variety of clothing, small targets, and algorithm architecture, the actual accuracy of the current clothing detection algorithm is low, the speed is slow, the generalization is weak, and the user experience is poor. Summary of the invention

[0009] In view of this, the present invention aims to overcome the shortcomings of the above-mentioned problems in the prior art and proposes a lightweight real-time clothing detection method for advanced manufacturing industry.

[0010] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0011] A first aspect of the present invention provides a lightweight real-time clothing detection method for advanced manufacturing, comprising the following steps:

[0012] Step 1: Collect clothing data images and build a clothing image dataset;

[0013] Step 2: Perform data augmentation and preprocessing on the collected data;

[0014] Step 3: Take YOLOv10 as the baseline model, improve YOLOv10 and build a clothing detection model;

[0015] Step 4: Send the data to the clothing detection model for training and testing;

[0016] Step 5: Use the trained clothing detection model to perform clothing detection.

[0017] Furthermore, in step 1, clothing image data of personnel in advanced manufacturing production scenarios are collected to construct a complete data set, and the data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0018] Furthermore, in step 2, the collected clothing data images are preprocessed by random scaling, image flipping, image brightness adjustment, and Mosaic data enhancement, and the diversity of training data samples is increased and the small sample data is expanded.

[0019] Furthermore, in step 3, the YOLOv10 network is improved by introducing the lightweight SOTA network MobileNetV4 and the attention mechanism Coordinate Attention, and adding the dynamic convolution ODConv and the small target detection head to construct a clothing detection model.

[0020] Furthermore, it also includes inserting a DA attention mechanism before the small target detection head. The DA attention mechanism first uses a dilated convolution, performs average pooling in the X and Y dimensions after the dilated convolution processing, and then merges the feature maps. Two convolution layers are used to splice the model input and the output after the feature map is merged, and after sigmoid function processing, it is sent to subsequent processing.

[0021] Furthermore, it also includes evaluating the clothing detection model using precision, recall, mAP, FPS, and model size.

[0022] A second aspect of the present invention provides a lightweight real-time clothing detection device for advanced manufacturing industry, comprising the following steps:

[0023] A data collection unit, used to collect clothing data images and construct a clothing image data set;

[0024] A data preprocessing unit, used to perform data enhancement and preprocessing on the collected data;

[0025] A model building unit, which is used to improve YOLOv10 and build a clothing detection model using YOLOv10 as the baseline model;

[0026] Model training unit, used to input data into clothing detection model for training and testing;

[0027] The clothing detection unit is used to perform clothing detection using a trained clothing detection model.

[0028] A third aspect of the present invention provides an electronic device, comprising a processor and a memory that is communicatively connected to the processor and is used to store executable instructions of the processor, wherein the processor is used to execute the above-mentioned lightweight real-time clothing detection method for advanced manufacturing industry.

[0029] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned lightweight real-time clothing detection method for advanced manufacturing.

[0030] Compared with the prior art, the lightweight real-time clothing detection method for advanced manufacturing industry described in the present invention has the following advantages:

[0031] According to the current work dress code specifications of major advanced manufacturing industries and their dress type requirements, the present invention collects and produces a dress data set containing mainstream dress types and different styles. The diverse data improves the generalization of the algorithm.

[0032] The present invention uses a lightweight backbone network to replace the original algorithm network, which greatly reduces the parameterization of the backbone network and improves the algorithm speed.

[0033] The present invention uses dynamic convolution and CA attention mechanism to ensure the accuracy of algorithm feature expression and help the algorithm further improve algorithm accuracy.

[0034] The present invention adds a special detection head for small targets and designs a special attention mechanism for the small target detection head to improve the detection accuracy of small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:

[0036] Figure 1 A flowchart of a lightweight real-time clothing detection method for advanced manufacturing industry is provided for the implementation of the present invention.

[0037] Figure 2 A baseline model block diagram of a lightweight real-time clothing detection algorithm for advanced manufacturing provided for the implementation of the present invention.

[0038] Figure 3 A backbone network replacement block diagram of a lightweight real-time clothing detection algorithm for advanced manufacturing industries provided for the implementation of the present invention.

[0039] Figure 4 A CA attention mechanism replacement block diagram of a lightweight real-time clothing detection algorithm for advanced manufacturing industries provided for the implementation of the present invention.

[0040] Figure 5 A block diagram of a DA attention mechanism for designing a lightweight real-time clothing detection algorithm for advanced manufacturing provided for the implementation of the present invention.

[0041] Figure 6 An overall block diagram of a clothing detection model provided for the implementation of the present invention. DETAILED DESCRIPTION

[0042] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0043] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0044] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.

[0045] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0046] Embodiment 1:

[0047] like Figure 1 As shown, the present invention provides a lightweight real-time clothing detection method for advanced manufacturing industry, comprising the following steps:

[0048] Step 1: Collect clothing data images and build a clothing image dataset.

[0049] Step 2: Preprocess the clothing data.

[0050] Step 3: Take YOLOv10 as the baseline model, improve YOLOv10, replace the backbone network with MobileNetV4, introduce ODConv and CA attention mechanism, add a new small target detection head and introduce DA attention mechanism, and build a clothing detection model.

[0051] Step 4: Feed the data into the improved YOLOv10 network for training.

[0052] Step 5: Test the test data set and compare the test results of the improved model with the test results of the mainstream algorithm for verification.

[0053] In step 1 of the present invention, according to the current work dress code of the advanced manufacturing industry and the requirements for the types of dress, the clothing of personnel in the production environment is classified and determined, with a total of 12 categories of data. They are laboratory coats, isolation gowns, tight-fitting protective clothing, one-piece protective clothing, masks, single-piece hair covers, hair covers, latex gloves, plastic gloves, hands, goggles and protective masks. According to these 12 categories, data collection and labeling are performed, and the data set is divided into a training set, a validation set and a test set in a ratio of 8:1:1.

[0054] In step 2 of the present invention, data enhancement and preprocessing are performed on the collected data. For the collected data on clothing, random scaling, image flipping, image brightness adjustment, Mosaic data enhancement and other methods are used to preprocess the images to increase the diversity of training data samples. For categories with a small number of samples, methods such as copying and splicing are used to expand the small sample data.

[0055] like Figure 2 The figure shows the YOLOv10 baseline network structure. In the present invention, the lightweight real-time detection network replaces the backbone network of YOLOv10 with the MobileNetV4 network. The MobileNetV4 network is a new generation of lightweight SOTA classification network, which introduces the universal bottleneck search block (UIB) and proposes an attention module tailored for mobile accelerators, Mobile MQA. Replacing the backbone network with MobileNetV4 not only reduces the number of parameters but also retains the extracted feature information to the greatest extent. The UIB module in the backbone network is divided into the FusedIB module, the IB module, the ExtraDW module and the ConvNeXt module according to the setting of the depthwise separable convolution, as shown in FIG. Figure 3 Shown is a backbone network replacement block diagram provided by an embodiment of the present invention.

[0056] Figure 4 A block diagram of the CA attention mechanism replacement provided for an embodiment of the present invention. At the end of the backbone model of YOLOv10, the CA (Coordinate Attention) attention mechanism is introduced to replace the previous PSA local self-attention module. The CA attention module not only considers channel information, but also direction-related position information, which can improve the accuracy of feature expression. Moreover, the CA attention module is small in size and can be effectively inserted into the network model, which improves the accuracy of feature expression while reducing the number of model parameters.

[0057] Specifically, CA attention is divided into two steps: Coordinate information embedding and Coordinate attention generation.

[0058] The method of embedding coordinate information is to aggregate features along two spatial directions respectively to obtain a pair of direction-aware feature maps. The calculation formula is as follows:

[0059]

[0060] Coordinate attention generation is done by concat operation after the transformation in information embedding, and then the convolution transformation function is used to transform it. The calculation formula is as follows:

[0061] f=δ(F1([z h ,z w ]))

[0062] Where δ is the sigmoid activation function. To reduce the complexity and computational overhead of the model, an appropriate reduction ratio is used to reduce the number of channels, and then the output and are expanded. The final output calculation formula of Coordinate Attention is as follows:

[0063]

[0064] Figure 5 Block diagram of the DA attention mechanism provided for an embodiment of the present invention. Often in the production environment of advanced manufacturing, protective measures such as masks and goggles are usually required. In actual detection images, they are often very small. For this type of small target, a new detection head is added on the basis of the improved YOLOv10, which is specifically used to detect smaller targets of size less than 32×32. This detection head is different from other detection heads. An attention mechanism DA for small target detection is designed and inserted before the small target detection head. This attention mechanism has a certain improvement on small targets and can improve the detection efficiency of small targets. The DA attention mechanism first uses a dilated convolution, and after the dilated convolution processing, average pooling is performed in the X and Y directions, and then the feature maps are merged. Then two convolution layers are used to splice the output after the model input and the feature map are merged, and after the sigmoid function processing, it is sent to the subsequent processing module.

[0065] Figure 6 The overall structural block diagram of the improved YOLOv10 network provided for the embodiment of the present invention. While reducing the parameters, the accuracy of the model may be affected. Therefore, in order to improve the detection accuracy of the model for clothing in the advanced manufacturing production environment, it is decided to introduce dynamic convolution ODConv before the detection head. ODConv introduces a multi-dimensional attention mechanism. It not only learns and applies attention weights for the number of convolution kernels, but also learns and applies attention weights for the spatial size, input channels and output channels of each convolution kernel. It ensures a more comprehensive and fine-grained dynamic adjustment of the convolution kernel based on the input features. And by utilizing a more detailed and diverse attention mechanism, ODConv can achieve better performance with fewer parameters without significantly increasing the size of the model.

[0066] In the present invention, the constructed training data set is sent to the improved YOLOv10 algorithm model of the present invention for training. The algorithm environment is configured using Anaconda, and the environment requirements of YOLOv10 are still the same. The specific environment requirements are as follows: torch==2.0.1, torchvision==0.15.2, onnx==1.14.0, onnxruntime==1.15.1, pycocotools==2.0.7, PyYAML==6.0.1, scipy==1.13.0, onnxslim==0.1.31, onnxruntime-gpu==1.18.0, gradio==4.31.5, op encv-python==4.9.0.80,psutil==5.9.8,py-cpuinfo==9.0.0,huggingface-hub==0.23.2,safetensors==0.4.3,matplotlib>=3.2.2,numpy>=1.18.5,Pillow>=7.1.2,requests>=2.23.0,thop>=0.1.1,tqdm>=4.64.0,pandas>=1.1.4,seaborn>=0.11.0。 Training hardware environment: GPU: NVIDIA GeForce RTX 2080Ti×8, video memory: 11G×8. Memory: 125G。 In the data model parameter setting and training parameter setting, select the corresponding YOLOv10 network structure configuration file and the pre-trained model corresponding to the configuration file. The training parameters are set as follows: learning rate lr = 0.001, number of iterations epoch = 300, batch size batchsize = 16, image input size imgsz = 640, momentum parameter momentum = 0.937, optimizer weight decay parameter weight_decay = 0.0005, non-maximum suppression threshold nms_threshold = 0.25, confidence parameter threshold conf_threshold = 0.45.

[0067] The present invention tests the trained YOLOv5 model, the native YOLOv10 model and the improved YOLOv10 model on the test set, and compares the experimental results. The evaluation indicators are precision, recall, mAP, FPS and model size, which verifies the advantages of the algorithm of the present invention in clothing detection accuracy and efficiency in advanced manufacturing production scenarios.

[0068] The test results are shown in the table below.

[0069] Table 1 Comparison of test results

[0070]

[0071]

[0072] As can be seen from the above table, the algorithm accuracy of the present invention is almost the same as that of the original YOLOv10, and it is 3.1% higher than that of YOLOv5. In terms of FPS, the present invention has a significant improvement over the original YOLOv10 and a significant improvement over YOLOv5. In terms of model size, the present invention is about half of the original YOLOv10 and about one-third of the YOLOv5 model.

[0073] It can be seen from the experimental results that the accuracy of the model of the present invention is comparable to that of the native YOLOv10, the model size is reduced, and the FPS is significantly improved, which fully meets the requirements of real-time detection in the production environment of advanced manufacturing.

[0074] In summary, the present invention provides a lightweight real-time clothing detection method for advanced manufacturing, which can better adapt to different types of advanced manufacturing production environments. The use of a lightweight backbone network and attention mechanism improves the detection accuracy and real-time performance of the algorithm, greatly improves the recognition accuracy and user experience, and has important production practice significance. The new model structure of the present invention enables the algorithm to perform clothing detection on personnel in the production environment in real time while maintaining detection accuracy. It also optimizes the detection method for small targets, improves the accuracy and robustness of the algorithm for small target detection, and thus improves the accuracy of the algorithm.

[0075] Embodiment 2:

[0076] A lightweight real-time clothing detection device for advanced manufacturing industry comprises the following steps:

[0077] A data collection unit, used to collect clothing data images and construct a clothing image data set;

[0078] A data preprocessing unit, used to perform data enhancement and preprocessing on the collected data;

[0079] A model building unit, which is used to improve YOLOv10 and build a clothing detection model using YOLOv10 as the baseline model;

[0080] Model training unit, used to input data into clothing detection model for training and testing;

[0081] The clothing detection unit is used to perform clothing detection using a trained clothing detection model.

[0082] Embodiment three:

[0083] An electronic device comprises a processor and a memory which is in communication with the processor and is used to store instructions executable by the processor. The processor is used to execute the above-mentioned lightweight real-time clothing detection method for advanced manufacturing industry.

[0084] Embodiment 4:

[0085] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned lightweight real-time clothing detection method for advanced manufacturing industry.

[0086] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A lightweight real-time clothing detection method for advanced manufacturing, characterized by: The steps include: Step 1: Collect clothing data images and build a clothing image dataset; Step 2: Perform data augmentation and preprocessing on the collected data; Step 3: Take YOLOv10 as the baseline model, improve YOLOv10 and build a clothing detection model; Step 4: Send the data to the clothing detection model for training and testing; Step 5: Use the trained clothing detection model to perform clothing detection.

2. A lightweight real-time clothing detection method for advanced manufacturing industry according to claim 1, characterized in that: In the step 1, clothing image data of personnel in advanced manufacturing production scenes are collected to construct a complete data set, and the data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:

1.

3. A lightweight real-time clothing detection method for advanced manufacturing industry according to claim 1, characterized in that: In the step 2, the collected personnel clothing data images are preprocessed by random scaling, image flipping, image brightness adjustment, and Mosaic data enhancement, and the diversity of training data samples is increased and the small sample data is expanded.

4. The lightweight real-time clothing detection method for advanced manufacturing industry according to claim 1, characterized in that: In step 3, the YOLOv10 network is improved by introducing the lightweight SOTA network MobileNetV4 and the attention mechanism Coordinate Attention, and adding the dynamic convolution ODConv and the small target detection head to build a clothing detection model.

5. A lightweight real-time clothing detection method for advanced manufacturing industry according to claim 4, characterized in that: The method also includes inserting a DA attention mechanism before the small target detection head. The DA attention mechanism first uses a dilated convolution, performs average pooling in the X and Y dimensions after the dilated convolution processing, merges the feature maps, and uses two convolution layers to concatenate the model input and the output after the feature map is merged. After the sigmoid function processing, it is sent to subsequent processing.

6. The lightweight real-time clothing detection method for advanced manufacturing industry according to claim 1, characterized in that: It also includes evaluating the clothing detection model using precision, recall, mAP, FPS, and model size.

7. A lightweight real-time clothing detection device for advanced manufacturing, characterized in that: The steps include: A data collection unit, used to collect clothing data images and construct a clothing image data set; A data preprocessing unit, used to perform data enhancement and preprocessing on the collected data; A model building unit, which is used to improve YOLOv10 and build a clothing detection model using YOLOv10 as the baseline model; Model training unit, used to input data into clothing detection model for training and testing; The clothing detection unit is used to perform clothing detection using a trained clothing detection model.

8. An electronic device, comprising a processor and a memory connected to the processor for storing instructions executable by the processor, characterized in that: The processor is used to execute a lightweight real-time clothing detection method for advanced manufacturing industry as described in any one of claims 1-6 above.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the computer program implements a lightweight real-time clothing detection method for advanced manufacturing industries as described in any one of claims 1 to 6.